93 Commits
Author SHA1 Message Date
patrick 0220f2d970 feat: fetch the next page of books before the reader hits the bottom
The prefetch zone is a viewport deep rather than 200px, and the observer
re-observes after a page lands so a chain of loads is not cut short.
2026-08-17 13:48:40 -04:00
patrick 55e00ba960 feat: import a Calibre library
Reads metadata.db and copies the books into a library — from a zip uploaded on
the library settings page, or from a path with `litestar calibre-import`. The
source is never touched, and re-running only picks up what is new.

Also names the formats mimetypes does not know: a Calibre library is full of
MOBI and AZW3, and a null content type used to fail the book endpoint.
2026-08-17 13:38:44 -04:00
patrick 85367daf7e feat: show books as an import adds them, not once it ends
Refresh the list every couple of seconds while the queue runs, holding off
while the reader has scrolled past the first page.
2026-08-17 11:21:49 -04:00
patrick 7e33a8fe05 Merge branch 'duplicate-detection' 2026-08-17 11:06:23 -04:00
patrick 202cbed30e refactor: name the merge options and drop their subtext 2026-08-17 11:06:19 -04:00
patrick a3443d36f1 feat: offer a merge that keeps the first book, without the workbench 2026-08-17 10:35:35 -04:00
patrick 75fe41e266 fix: reopen the merge dialog after dismissing it 2026-08-17 10:35:24 -04:00
patrick d16a90f08f feat: draw generated covers where a book has none 2026-08-17 10:35:01 -04:00
patrick 904d8fc76f refactor: move duplicates into library settings 2026-08-17 10:32:57 -04:00
patrick 315149e8a2 docs: brief for moving duplicates into library settings
Option B — libraries expand in the settings nav, duplicates becomes a section
within a library's pane.
2026-08-16 20:21:25 -04:00
patrick 699a1a7fa2 feat: merge books from the duplicates screen and the toolbar
A workbench with the folded records in a rail, the survivor's fields live and
editable beside them, and per-field actions chosen by what kind of field it is.
Fields the records agree on stay out of the way. Reachable from a duplicate
group and from the selection toolbar, which is the only way to merge a pair the
detector never proposed.
2026-08-16 20:21:19 -04:00
patrick 22539644a9 feat: merge books into one record
Files move on disk before any row moves, then everything repoints in one
transaction and the folded records are deleted. Progress keeps whichever is
furthest, shelves and tags union, and identifiers move only under a name the
survivor lacks. Metadata is left alone unless the caller resolves it, since
choosing between two titles is a judgement this endpoint cannot make. Nothing
is removed from disk.
2026-08-16 20:21:10 -04:00
patrick a48be517e4 fix: type the ids filter from its configured id type
advanced_alchemy annotates the `ids` query parameter as list[str] whatever the
config says, so a bigint primary key was compared against strings and Postgres
refused. Nothing had called ?ids= until now.
2026-08-16 20:20:44 -04:00
patrick 2e0d556c33 fix: return the publisher an EPUB declares
The lookup discarded its own result and fell off the end of the function, so
no EPUB ever contributed a publisher.
2026-08-16 00:01:02 -04:00
patrick ff75f2c758 chore: clear the lint in the files this branch touches
Unused imports, duplicated import lines, and bare excepts that swallowed
KeyboardInterrupt along with everything else.
2026-08-15 22:01:04 -04:00
patrick e967019964 feat: split the edition out of a book's title
"Fluent Python, 2nd Edition" is one book with a field for the edition. Left in
the title it also splits the library, since the second edition never looks like
the first. Only a numbered statement is moved, so "Catch 22" keeps its number
and "Global Edition" — which has nowhere to go in an integer column — stays put.
2026-08-15 21:57:17 -04:00
patrick f6bb06ac6e fix: merge identifiers across a book's formats
Identifiers are a collection, but the whole dict was replaced per format, so
the last file to report won outright — an EPUB declaring an ASIN, a Google id
and a Calibre id kept none of them once a PDF contributed one ISBN. Accumulate
instead, letting a declared identifier outrank one scraped off a page.
2026-08-15 21:57:01 -04:00
patrick 8ee406533c feat: canonicalise author names
Extractors wrote whatever the file said, so one person held several rows:
"Sam Newman" beside "Newman, Sam;" beside "Sam Newman.epub", the last because
the upload path never stripped the file extension. Tidy on write, in the
validator and in the uniqueness lookup alike, and merge the rows that collide.
2026-08-15 21:56:35 -04:00
patrick 373c96d6d6 fix: match the compact edition markers a cover carries
"Building Microservices, 2E" never matched "Building Microservices". Key the
strip on the trailing "e" so 2E, 5e and 3 Ed are caught, while a bare number
leaves "Catch 22" and "Blade Runner 2049" alone. Stored keys are recomputed.
2026-08-15 21:55:45 -04:00
patrick fc6b97bf38 feat: review and dismiss duplicate books
Adds the per-library review screen, a muted note in the upload tray for a book
that may already be held, and the remote functions behind them. The tray note
is deliberately weaker than the file-level one, which means something stronger.
2026-08-15 21:55:20 -04:00
patrick 5047277845 feat: detect books that may already be in the library
Match on a shared identifier, or on a normalised title credited to a shared
author, and report candidates rather than refusing anything — a metadata match
is a guess, and a second edition is not a mistake. Adds a library-wide review
pass, dismissals, and renames the fingerprint pre-flight to duplicate-files.
2026-08-15 21:55:13 -04:00
patrick 6d1890ce04 feat: store matching keys and duplicate dismissals
Derive normalized_title, normalized_name and normalized_value with @validates
so no write can bypass them, and add the table recording pairs a reader has
said are not the same book.
2026-08-15 21:54:32 -04:00
patrick 523117ec28 fix: keep the identifiers an EPUB declares
DC:identifier was validated verbatim, so hyphenated and urn:isbn: forms never
reached the checksum and every non-ISBN identifier was discarded. Normalise
first, and name whatever survives.
2026-08-15 21:54:12 -04:00
patrick 86e1d096ef feat: add normalisation helpers for book matching
Reduce a title, an author and an identifier to a single comparison key, so
two copies of one book can be recognised by equality rather than by a
similarity score.
2026-08-15 21:44:50 -04:00
patrick 3f39f1f8ae fix: drop the add-anyway override from the upload tray
Storing the same bytes twice splits reading progress and shelves across two
records that can never converge. The skipped row still names the book that
holds the file; allow_duplicates stays on the API for a wrong verdict.
2026-08-14 14:18:08 -04:00
patrick d321315acf fix: ignore duplicate matches whose file is gone
The hash lives in the database and the file does not, so a file deleted behind
the app's back went on refusing its own replacement. Matches are now checked
against disk, and re-adding a book's own missing file writes it back into the
row that already describes it.
2026-08-14 14:18:08 -04:00
patrick b124a65d6e feat: report skipped duplicates in the upload tray
A book whose files were all already stored settles as skipped rather than done,
naming the book that holds them and offering to add it anyway.
2026-08-13 17:23:57 -04:00
patrick d78b21c27f feat: check for duplicate files when importing books
Incoming files are matched against what is already stored, keyed on the
KOReader hash and the file size. Bulk uploads skip and report them, deliberate
creates are refused with a 409, the consume directory parks them aside, and
allow_duplicates overrides all three.

Also: books whose metadata generates a path another book already owns are moved
aside, so a forced copy cannot overwrite the original's files.
2026-08-13 17:23:48 -04:00
patrick 968166c1fd Merge branch 'ui-rewrite' 2026-08-13 13:52:47 -04:00
patrick 8589adbd1b docs: record the outstanding backend issues in TODO.md 2026-08-13 00:42:54 -04:00
patrick 510306f24d fix: keep embedded metadata and group formats on directory upload
The filepath extractor was merged on the right, so a parent folder's name
overrode the title inside the file. Grouping also split a book's formats when
its own folder was the upload root.
2026-08-12 15:17:48 -04:00
patrick bd8d68b9ba fix: resolve file paths and stream the multi-book zip download
file.path is relative to book.path, so the archive matched nothing on disk and
came out empty. Also namespace entries per book to stop filename collisions, and
stream the zip instead of buffering it whole.
2026-08-12 15:17:48 -04:00
patrick 540522e828 fix: reuse existing link rows when updating book relationships
Assigning through the association proxy recreated links for targets the book
already had, colliding with the link tables' unique constraints. Reconcile the
collections in place instead, identifiers included.
2026-08-12 15:17:48 -04:00
patrick 5305d3bb5e docs: correct the CSP entry in TODO.md
Scripted EPUBs are a regression from the foliate migration, not a pre-existing
gap: epub.js sandboxed without allow-scripts by default, foliate sets it
unconditionally. Records grimmory's fix — CSP on per-entry responses rather
than app-wide — which also sidesteps the mode-watcher blocker, plus the
whole-file buffering it would remove.
2026-08-12 14:21:10 -04:00
patrick 7e04826fa5 feat: draw covers for books without one
Full-bleed serif title on a woven ground, drawn in the browser rather than
stored, sized with container queries so one component works from 36px to full
page. Colour is hashed into a list of bookcloth tones instead of the hue wheel.

The grid and search were falling back to default_cover.jpg because a missing
cover made the src "/api/undefined"; they now check for one first.
2026-08-12 13:44:14 -04:00
patrick 96789620bb feat: run book uploads in a background tray
One request per book instead of one for everything, queued in state above the
dialog so closing it no longer cancels the import. A docked tray reports each
book and offers a retry; it clears itself after a clean run, stays put if
anything failed, and holds while hovered.

Also fixes the dialog overflowing on long filenames — Dialog.Content is a grid
and its body needed min-w-0 to shrink below the content's intrinsic width.
2026-08-12 13:42:40 -04:00
patrick d4bdb5ed42 feat: accept files and folders when uploading books
The picker could only choose folders: webkitdirectory switches the OS dialog
into folder mode rather than filtering, so one input cannot offer both. Adds
a drop zone with two browse buttons behind two inputs.

Dropping a folder never worked either — dataTransfer.files does not descend
into directories, so it arrived as one typeless entry and was rejected. Walks
dataTransfer.items instead.

Also collects skipped files into one expandable line rather than a toast each,
which a folder of covers and notes made unusable, and shows the file count and
total size before uploading.
2026-08-12 12:04:14 -04:00
patrick 5f2d68694d feat: rebuild the edit dialog around a cover and files rail
Drops the three tabs for a wide dialog: cover and files in a fixed rail,
all fourteen metadata fields grouped beside them, Save in the footer.

Adds file management, which the Files tab never had — it could only upload,
never list or remove. Files now show size and format with download and
delete, the latter asking whether to remove it from disk too.
2026-08-12 11:40:49 -04:00
patrick d6207b5743 refactor: resolve() internal links instead of plain hrefs
Type-checks every link against the real route tree. Caught a dead one: the
book page linked tags to /tag/{id}, a route that has never existed.
2026-08-12 11:08:09 -04:00
patrick 51c31e6bf6 docs: document the foliate-js reader
Adds a reader section covering the vendored tree, the $foliate alias and the
traps around it, and refreshes the stale stack line, oklch claim and rough
edges. Records why CSP is not enabled in TODO.md.
2026-08-12 01:37:19 -04:00
patrick 961a63480e chore: drop epubjs
Nothing imports it since the foliate rewrite. Also note the one-shot purge of
its locations cache in TODO.md, so the temporary code gets removed later.
2026-08-11 23:15:39 -04:00
patrick 92ffa4f7c2 fix: forward app shortcuts pressed inside the book
Key events do not cross document boundaries, so shortcuts bound on the window
never fired while the book had focus and the browser default won instead —
ctrl+B opened bookmarks rather than the chapter sidebar. Replay them on the
window, cancelling the original only if a handler claimed it so ctrl+C still
copies.
2026-08-11 23:12:00 -04:00
patrick 9155129ccd fix: make the reader's typography, width and arrow keys work
Typography was in the stylesheet foliate prepends, which the book overrides;
moved to the appended one. The window keydown handler was never bound.

Also: wider reading area, page on a card, arrows moved onto it, spread gutter,
and the sidebar stays open on chapter clicks.
2026-08-11 22:45:14 -04:00
patrick 6318115380 fix: make the vendored slider track visible
shadcn-svelte generates the track with `data-horizontal:` / `data-vertical:`
variants, which Tailwind compiles to `[data-horizontal]` — an attribute
nothing sets, since the orientation is carried as data-orientation. The track
got no height, leaving only the thumb on screen. Use the `data-[orientation=…]`
form the rest of ui/ already uses.
2026-08-11 22:43:22 -04:00
patrick fe12f4be76 feat: add the reader settings panel
Typeface, size, weight, line height, letter spacing, justify and hyphenate,
plus flow, columns, line width, gap and margins — persisted in localStorage
and applied live.

Fixed-layout books hide the typography and column controls: they render
through foliate-fxl, which has no setStyles and ignores the paginator's
attributes, so those sliders would do nothing.
2026-08-11 22:27:04 -04:00
patrick 190e7af76d feat: replace the epub reader with foliate-js
Drops epub.js's 1600-location pre-pass and its localStorage cache: foliate
hands back a CFI and an overall fraction on every relocate, so first open no
longer freezes.

Also fixes resuming (pre-resolves the stored CFI, falling back to the stored
percentage rather than silently resetting to page one), progress saving
(missing Content-Type, and no flush on the way out), and arrow keys inside
the book iframe.
2026-08-11 22:21:15 -04:00
patrick 783f4d226d feat: add reader settings state and stylesheet generator
Settings are validated on read, not only on write: several go straight onto
foliate's renderer as custom-element attributes, where a bad value wedges the
layout silently rather than throwing. Unknown keys are stripped and missing
ones filled from the defaults, so settings added later do not invalidate what
a reader already has stored.

The stylesheet is built as foliate's [before, after] pair. The paginator
prepends the first to the section's <head> and appends the second, so the
book's own CSS overrides the first and loses to the second. Typography goes
in the first — a book that styles its own headings still gets to, and font
size lands on <html> without !important so rem/em headings scale rather than
being flattened. Colour goes in the second, because most EPUBs set a body
background and would otherwise leave the page white against a dark UI.

Font choices are reused from the app theme's FONT_STACKS rather than
duplicated, so the reader cannot drift from the rest of the UI.

Rename the ambient declaration to foliate-js.d.ts: as foliate.d.ts beside
foliate.ts, TypeScript takes it for that file's emitted declaration, drops it
from the program, and every $foliate import fails with TS2307.
2026-08-11 22:08:58 -04:00
patrick c946b6d71c chore: wire vendored foliate-js into the build
$foliate is a Vite-only alias, deliberately absent from kit.alias and
tsconfig paths: TypeScript cannot resolve it, so the ambient declaration in
src/lib/reader/foliate.d.ts is the only candidate and svelte-check never
walks the ~11k lines of untyped vendored JS. The generated tsconfig includes
src/**/*.js and checkJs is on, so src/lib/vendor is excluded as well —
`exclude` replaces rather than merges, hence the repeated service-worker
entries.

optimizeDeps.include for construct-style-sheets-polyfill: it sits behind a
dynamic import in fixed-layout.js, so Vite would otherwise discover it
mid-session and force a page reload on the first fixed-layout EPUB.

Verified with a temporary import that the graph resolves: view.js and
paginator.js are emitted as chunks, and the bundled pdf.js is our stub
rather than upstream's bare `import '@pdfjs/pdf.min.mjs'`.
2026-08-11 22:04:50 -04:00
patrick dd65e34869 chore: vendor foliate-js
foliate-js has no npm release and no build step; upstream recommends a git
submodule. Copy it instead: this repo has no submodules and already vendors
pdf.js the same way under static/pdfjs, and a submodule would pull in 231
files / 13 MB of which 191 files / 12 MB is a bundled pdf.js build that
Chitai does not use.

Only the import closure reachable from view.js is vendored — 15 files, 584K.
pdf.js is replaced by a stub that throws: view.js reaches it through a
static-string dynamic import inside makeBook, which Rollup resolves at build
time even though Chitai serves PDFs from static/pdfjs/web/viewer.html, and
upstream's version opens with a bare `import '@pdfjs/pdf.min.mjs'` that does
not resolve here.

scripts/vendor-foliate.sh pins the commit and makes the next update a one-line
change. fixed-layout.js needs construct-style-sheets-polyfill, so add it.
2026-08-11 21:59:33 -04:00
patrick a1281f129c chore: unbreak prettier by bumping prettier-plugin-tailwindcss
prettier-plugin-tailwindcss 0.6.14 is incompatible with prettier 3.8.1 and
threw "getVisitorKeys is not a function" on every .svelte file, so `pnpm
lint` could not run at all. 0.8.1 fixes it.

No reformatting here: the files prettier now reports are mostly vendored
shadcn components, and rewriting those is unrelated to any current work.
2026-08-11 21:54:17 -04:00
patrick ee46f5568c fix: resume epub progress from the fields the API returns
The load function read progress.epub_loc and progress.progress, neither of
which BookProgressRead carries, so resuming never worked and every open
started at page one. Read epub_cfi and percentage instead.

Fail loudly when the book cannot be fetched. The previous version caught
everything and returned {status, error} that nothing consumed, so a missing
book produced a reader sitting on its spinner forever. Check response.ok
before parsing too: fetch resolves on a 404, and arrayBuffer() happily
returns the error page, which only surfaced later as an opaque epub parse
error.

Give the reader a working sidebar, a theme that survives the book's own
CSS, and an error state with a retry. The theme is registered rather than
applied as bare overrides because most EPUBs ship their own body colours
and win otherwise, which is why the page stayed white against a dark UI.
2026-08-11 21:47:12 -04:00
patrick ac5a5c75aa fix: surface reader load failures instead of spinning forever
The PDF reader pointed an iframe at the vendored pdf.js viewer and had no
way to report a failure: the viewer renders its own errors inside the
frame, so a missing or unreadable file showed an empty grey panel. Probe
the file with a request first and render an error card with a retry when
that fails, matching the chrome the EPUB reader uses.

Add the missing +page.ts for the PDF route so the tab gets the book title
rather than a bare "Reader".

Drop chapter-sidebar.svelte and its import. It rendered a list of chapters
whose buttons had no click handler, and read/+layout@.svelte imported it
without ever rendering it. The reader's own sidebar replaces it.
2026-08-11 21:47:03 -04:00
patrick 49f94f9ee1 fix: mount ModeWatcher and Toaster at the root layout
The reader and login pages use +layout@ breakouts to escape the (root)
shell, so anything mounted in (root)/+layout.svelte never reaches them.
ModeWatcher was mounted there and duplicated in login/+layout@.svelte,
which left the reader's theme toggle inert: it flipped the store but
nothing wrote .dark onto <html>.

Mount both once at the top-level +layout.svelte instead, which every
route inherits regardless of breakouts.
2026-08-11 21:46:50 -04:00
patrick 2aba90f910 chore: clear epubjs type import and empty server load 2026-08-11 20:23:31 -04:00
patrick d2892353cb fix: reset the edit dialog when switching books 2026-08-11 20:23:31 -04:00
patrick d8f8a0ab95 docs: add TODO.md 2026-08-11 20:08:53 -04:00
patrick 4ee127b6cb fix: remove duplicate separator in the shelves dropdown 2026-08-11 20:08:53 -04:00
patrick 6366582498 fix: silence state_referenced_locally in vendored ui components 2026-08-11 20:08:53 -04:00
patrick ef8e5e7fba fix: stop icons remounting and shelves collapsing into nothing 2026-08-11 20:08:53 -04:00
patrick c09b8365fa feat: put the library view mode in the url via shallow routing 2026-08-11 20:08:53 -04:00
patrick 0139f6f5eb feat: extract built-in view presets and add filter chips 2026-08-11 20:08:53 -04:00
patrick 4c3fd66a56 feat: rebuild the book page with the jacket layout 2026-08-11 20:08:53 -04:00
patrick f2b8f337f7 feat: rebuild the table view with sortable, toggleable columns 2026-08-11 20:08:53 -04:00
patrick 75360ff603 feat: add book cover component and detail-card list view 2026-08-11 20:08:53 -04:00
patrick 87fff20d72 feat: widen search and label the upload button 2026-08-11 20:08:53 -04:00
patrick adddcfeeed feat: add account menu to the sidebar footer 2026-08-11 20:08:53 -04:00
patrick cb070336c6 feat: replace the sidebar logo with a book mark 2026-08-11 20:08:53 -04:00
patrick f68a233dca refactor: replace hardcoded colours with semantic tokens 2026-08-11 20:08:53 -04:00
patrick f36aa14463 feat: add theme editor in settings 2026-08-11 20:08:53 -04:00
patrick d71b53b3c5 feat: add themeable design tokens with cookie-backed SSR 2026-08-11 20:08:53 -04:00
patrick a1f39a8dc8 fix: scope book progress and shelves to the requesting user 2026-08-11 20:08:26 -04:00
patrick dfeb4c3267 chore: remove unused book files accordion component 2026-08-11 00:32:44 -04:00
patrick e4ecfad5dd chore: regenerate OpenAPI types 2026-08-11 00:27:56 -04:00
patrick 3a3957d432 fix: resolve BookProgressRead reference in BookRead 2026-08-11 00:27:56 -04:00
patrick 23a12f7970 docs: add AGENTS.md for repo, backend and frontend 2026-08-10 23:58:42 -04:00
patrick c3e337de32 fix: correct book browser scroll area height 2026-08-10 23:58:42 -04:00
patrick 3ae3dcb0d0 fix: send correct field names when updating book progress 2026-08-10 23:58:42 -04:00
patrick 3091cc879d fix: skip unreadable PDFs during metadata extraction 2026-08-10 23:58:42 -04:00
patrick 1415cb6244 fix: preserve directory structure when uploading books 2026-08-10 23:58:42 -04:00
patrick cc39e79cd7 chore: regenerate initial migrations 2026-08-10 23:58:09 -04:00
patrick c5b703b75a refactor: use unique constraints instead of composite keys on link tables 2026-08-10 23:58:09 -04:00
patrick 9fe69641c5 feat: add library icons 2026-08-10 23:58:02 -04:00
patrick 0963ee85c2 chore: update nix dev shells and pnpm build settings 2026-08-10 23:58:02 -04:00
patrick 448e0e0090 chore: ignore local library files 2026-08-10 23:57:46 -04:00
patrick 9711c68fbb chore: add migration for initial db tables 2026-03-09 14:28:35 -04:00
patrick 20df5ea140 remove python313Full package as it is no longer available in nixos 25.11 2026-03-09 14:26:33 -04:00
patrick 2e1d16ef9e refactor: move db and file watcher setup to use async context manager 2026-03-09 14:25:24 -04:00
patrick 2de0aac23a chore: update dependencies and ignore new warning in tests
ignore JWT related warning that was introduced when updating
dependencies related to JWT secret size (test value of "secret" is too
short).
2026-03-09 14:21:37 -04:00
patrick 10eaf84329 refactor: simplify settings page layout 2026-03-09 14:19:59 -04:00
patrick a3e49dc918 regenerate OpenAPI schema with new KOSync models 2026-03-09 14:19:50 -04:00
patrick 3072f72958 feat: add device settings page
- Add device CRUD operations (create, delete, regenerate api keys)
- Add API key visibility toggle and copy button
2026-03-09 14:16:53 -04:00
patrick 51c1900d8c feat: add KOSync server
- Add KOSync device management
- Add API key auth middleware for devices to authenticate
- Add KOSync-compatible progress sync endpoints
- Add basic tests for KOSync compatible hashes
2026-03-09 14:11:21 -04:00
patrick 20a69de968 feat: add KOReader compatible hash to file metadata
Implement KOReader's partial MD5 algorithm for document identification. This hash allows KOReader devices to match local files with server records for reading progress synchronization (KOSync).
2026-03-09 13:38:07 -04:00
224 changed files with 33390 additions and 3048 deletions
+9
View File
@@ -5,6 +5,15 @@ CHITAI_TOKEN_SECRET=secret
CHITAI_DEFAULT_LIBRARY_NAME=Books
CHITAI_DEFAULT_LIBRARY_PATH="libraries/books"
# Duplicate detection when importing files (optional).
# Scope: "library" compares against the library being uploaded to, "global" against
# every library, "off" disables the check.
CHITAI_DUPLICATE_SCOPE=library
# Where the consume watcher parks files it refused as duplicates. Keep it outside
# CHITAI_CONSUME_PATH, or the watcher picks them straight back up.
CHITAI_DUPLICATE_PATH="duplicates"
# You probably should not change these
CHITAI_API_URL="http://backend:8000"
CHITAI_API_DEBUG=false
+3
View File
@@ -1,2 +1,5 @@
# Mark pdfjs as vendored code
frontend/static/pdfjs/** linguist-vendored
# Mark foliate-js as vendored code
frontend/src/lib/vendor/** linguist-vendored
+1
View File
@@ -1,3 +1,4 @@
.env
.postgres/
.venv
tmp/
+108
View File
@@ -0,0 +1,108 @@
# Chitai
Self-hosted eBook library manager. Users organise eBook files into **libraries**, which contain
**books** (one book = one metadata record + one or more files on disk). Books carry authors,
publishers, tags, series and identifiers, can be grouped into per-user **bookshelves**, and are read
in-browser through built-in EPUB and PDF readers with reading-progress tracking. The catalogue is
also exposed as an **OPDS** feed for e-reader apps, and progress syncs with **KOReader** devices via
a KOSync-compatible endpoint.
## Layout
| Path | What |
| -------------------------- | ---------------------------------------------------------------------------------------- |
| `backend/` | Litestar REST API + PostgreSQL. See `backend/AGENTS.md`. |
| `frontend/` | SvelteKit SSR web app. See `frontend/AGENTS.md`. |
| `frontend/src/lib/vendor/` | Vendored `foliate-js` (the EPUB engine), copied by `frontend/scripts/vendor-foliate.sh`. |
| `frontend/static/pdfjs/` | Vendored pdf.js viewer, used by the PDF reader in an iframe. |
| `docker-compose.yml` | Production stack: `db` (postgres:17), `backend`, `frontend`. |
| `docs/screenshots/` | Images used by `README.md`. |
| `shell.nix` | Root dev shell; composes the two sub-shells. |
## Development environment
`nix-shell` from the repo root is the intended entry point. It pulls in both
`backend/shell.nix` and `frontend/shell.nix`, whose `shellHook`s do real work as a side effect:
- **backend** — `uv venv` + `uv sync`, then `initdb` / `pg_ctl start` into `backend/.postgres/`
(socket dir, not TCP), `createdb chitai`, and finally applies migrations with
`alchemy --config chitai.database.config.config upgrade --no-prompt`.
`exitHook` stops PostgreSQL on shell exit.
- **frontend** — `pnpm install`.
So entering the shell gives you a running database with an up-to-date schema; you do not need to
start Postgres yourself. Both hooks `cd` around, which can be surprising in scripts.
Without Nix you need: Python 3.13 + `uv`, Node 24 + `pnpm`, and PostgreSQL 17 with the `pg_trgm`
extension available.
## Configuration
All backend settings are read by `backend/src/chitai/config.py` (`Settings`, pydantic-settings) with
the **`CHITAI_` prefix** from the repo-root `.env`. `.env.prod-example` is the template; copy it to
`.env` for a fresh deployment.
`.env` is gitignored and contains a real `CHITAI_TOKEN_SECRET` — do not print, copy or commit it.
The frontend reads exactly one variable, `VITE_BACKEND_API_URL`
(`frontend/src/lib/server/config.ts`, defaulting to `http://localhost:8000`).
## How a request flows
```
browser
└─ SvelteKit SSR node server
├─ hooks.server.ts reads `authToken` cookie, validates via GET /access/me,
│ populates locals.user / locals.api, redirects to /login otherwise
├─ $lib/api/*.remote.ts remote functions (query/command/form) → locals.api (ApiClient)
└─ routes/api/[...path] catch-all proxy, for browser-direct fetches (reader file streams)
└─ Litestar backend
└─ controllers/ → services/ → SQLAlchemy models → PostgreSQL
```
The JWT the frontend holds in the `authToken` cookie is the same bearer token the backend issues
from `POST /access/login`. The frontend never stores credentials beyond that cookie.
## Commands
Backend (from `backend/`):
```bash
uv run litestar --app-dir src/chitai/ run --reload # dev server on :8000
pytest tests/ # needs Docker (pytest-databases)
ruff format src/
alchemy --config chitai.database.config.config make-migrations
alchemy --config chitai.database.config.config upgrade
# Import a Calibre library. Copies files; --dry-run reports without writing.
litestar --app-dir src/chitai/ calibre-import <path> --library <slug>
```
Frontend (from `frontend/`):
```bash
pnpm dev # vite dev server on :5173
pnpm build # adapter-node output in build/
pnpm check # svelte-check — run before finishing
pnpm lint # prettier --check + eslint
pnpm format # prettier --write
```
API docs are served by the running backend at `http://localhost:8000/schema/` (Swagger) and
`/schema/openapi.json`.
## Cross-cutting rules
- **Keep the two schema layers in sync.** Backend request/response shapes live in
`backend/src/chitai/schemas/` (Pydantic); the frontend mirrors them as Zod schemas in
`frontend/src/lib/schema/`. Changing one without the other produces runtime validation failures,
not type errors.
- **Regenerate the OpenAPI types** after changing API shapes.
`frontend/src/lib/schema/openapi/schema.d.ts` is generated from the backend's OpenAPI document
with `openapi-typescript` (a devDependency; there is no `package.json` script for it, so it is run
manually against a live backend).
- **Migrations are mandatory.** The app runs with `create_all=False`, so a model change without a
matching Alembic revision will not reach the database.
- **Commit messages** follow `type: summary``feat:`, `fix:`, `refactor:`, `chore:`.
- The three `README.md` files are user-facing. Agent-facing knowledge belongs in the `AGENTS.md`
files.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+339
View File
@@ -0,0 +1,339 @@
# TODO
Known issues and deferred work. Agent-facing notes belong in the `AGENTS.md` files;
this is for things that are broken or missing and not yet scheduled.
## Backend
### Identifier extraction drops most ISBNs
`backend/src/chitai/services/metadata_extractor.py``EpubExtractor._extract_identifiers`
The EPUB path validates `DC:identifier` values verbatim:
```python
for id in epub.get_metadata("DC", "identifier"):
if is_valid_isbn(id[0]):
...
```
`is_valid_isbn` branches on `len(isbn)` being exactly 10 or 13, so anything carrying
formatting fails. Two consequences:
- **Hyphenated ISBNs are silently dropped.** `978-0-486-28211-4` is 17 characters, so it
never reaches the checksum. Most EPUBs write ISBNs hyphenated, so the majority are lost.
The PDF path already does this correctly — `_extract_isbns` calls
`match.replace("-", "")` before validating.
- **`urn:isbn:` prefixes are dropped** for the same reason. This is a common EPUB form.
Fix: normalise before validating — strip a leading `urn:isbn:`, then remove everything
that isn't `0-9` or `X`. Reuse the PDF path's approach rather than duplicating it.
Note this only runs at upload, so fixing it changes nothing for books already imported.
A backfill would need to re-read the files on disk.
### Non-ISBN identifiers are discarded
Same function. Anything that isn't a valid ISBN is thrown away, including values EPUBs
routinely carry: `urn:uuid:…`, `calibre:…`, Google Books volume IDs and ASINs.
The `Identifier` model is already generic (`name` + `value`, unique per book), so storing
them needs no schema change — only the extractor decides what survives. The frontend
already renders ISBN, ASIN and DOI as links and shows unknown types as plain values, so
anything stored will display sensibly.
Worth adding at the same time:
- A DOI regex (`10.\d{4,9}/\S+`) alongside the ISBN scan in `PdfExtractor._extract_isbns`
— academic PDFs carry one and it is the most useful identifier they have.
- Deriving ISBN-10 from ISBN-13 when only the latter is present. It is a pure checksum
conversion and doubles the chance of an external lookup matching.
### Any authenticated user can delete any library or book
`backend/src/chitai/database/models/library.py`, `backend/src/chitai/database/models/user.py`
There is no authorization tier. `Library` has no owner column, and `User` carries only
`email` and `password` — no role, no `is_active`. So every authenticated account can create
and delete libraries, and delete books along with their files on disk. Per-user scoping
exists only for reading progress and bookshelves, which `provide_book_service` restricts
correctly.
For a single-household deployment that may well be acceptable. The point is that it is
emergent rather than chosen. The cheapest meaningful step is an `is_admin` flag gating
library deletion and `delete_books` — a model change plus a migration.
### Basic auth answers a malformed header with a 500
`backend/src/chitai/middleware/basic_auth.py` — line 22
```python
username, password = b64decode(auth_header.split("Basic ")[1]).decode().split(":")
```
Nothing guards the parse. A `Bearer` token raises `IndexError`, non-base64 raises
`binascii.Error`, and a credential with no colon raises `ValueError` — as does a password
that *contains* one, since there is no `maxsplit=1`. Every case surfaces as a 500 on an
unauthenticated endpoint. All of them should be 401.
### An unknown KOSync API key returns 404
`backend/src/chitai/middleware/kosync_auth.py` — line 32
`KosyncDeviceService.get_by_api_key` uses `get_one`, which raises `NotFoundError`, but the
middleware catches only `PermissionDeniedException`. The global handler in
`exceptions/handlers.py` then renders it as a 404, so a device presenting a bad key is told
the route does not exist rather than that it is unauthorized. The same file still carries a
leftover `print(exc)`.
Worth doing at the same time: `KosyncDeviceService._generate_api_key` uses
`secrets.token_hex(8)`. 64 bits is thin for a long-lived bearer credential where 32 bytes
is the convention.
### The multi-book download cannot be driven from a test
`backend/src/chitai/services/book.py``BookService.get_files`
`/books/download` is the only handler returning a Litestar `Stream`, and it cannot be
exercised through `AsyncTestClient`. The request itself succeeds, then fixture teardown
hangs: the test transport never sends the `http.disconnect` that the streaming response
waits on, so the app's lifespan shutdown never completes. Coverage therefore sits at the
service level, on `get_files` directly.
Unresolved whether the endpoint also stalls behind a real ASGI server, where that
disconnect does arrive. Worth one manual check against `litestar run` before relying on it.
### The production image runs the development server
`backend/Dockerfile` — the final `CMD`
```
CMD ["litestar", "--app-dir", "chitai", "run", "--host", "0.0.0.0", "--port", "8000"]
```
`litestar run` is the CLI development runner. Production should invoke uvicorn or granian
directly, with a worker count.
### Nothing gates formatting, linting or types
`ruff format --check src/` reports 27 of 61 files unformatted, and `ruff check src/` finds
114 errors — 100 of them unused imports, the rest bare `except`, unused variables and
`== True` comparisons. `ruff check --fix` clears 51 automatically.
There is no `[tool.ruff]` section in `pyproject.toml`, so only ruff's default `E4/E7/E9/F`
rules run, and no type checker is configured at all despite `# type: ignore` comments in
the tree. Individually these are trivial; collectively they say nothing runs on commit.
## Frontend
### Scripted EPUBs run against the app origin
**This is a regression from the foliate-js migration, not a pre-existing gap.**
The old epub.js reader never passed `allowScriptedContent`. epub.js defaults it to
`false`, which sets `iframe.sandbox = "allow-same-origin"` — no `allow-scripts` — so
script inside a book never ran. The vendored foliate-js sets, unconditionally:
```js
// paginator.js — and the same in fixed-layout.js
// `allow-scripts` is needed for events because of WebKit bug
this.#iframe.setAttribute("sandbox", "allow-same-origin allow-scripts");
```
`allow-same-origin` together with `allow-scripts` is the combination that makes the
sandbox attribute do nothing. Sections are served as same-origin `blob:` URLs, so script
in a book can reach `/api/*` with the session cookie attached. foliate's README says as
much and tells you to use a CSP instead; we have not added one.
This is not theoretical. Audiobookshelf shipped the same combination and got
**CVE-2024-35236** — scripted EPUB plus an unrestricted upload gave remote code
execution; fixed in 2.10.0 by making scripted content a per-library opt-in, off by
default. Kavita (CVE-2024-39307) and Jellyfin (fixed 10.9.8) are variations on it.
Write-up: <https://gebir.ge/blog/every-trick-in-the-book/>.
**An app-wide CSP is the wrong shape.** `kit.csp` with `script-src: ['self']` also blocks
`mode-watcher`'s inline `setInitialMode`, which sets the dark class before first paint —
SvelteKit only nonces the bootstrap script it injects itself, so every page load would
flash the light theme. Pinning a hash of a third-party inline script breaks silently on
upgrade.
**Grimmory solves it properly**, and it runs foliate-js too. Rather than handing foliate
the whole file, it serves each EPUB entry from its own endpoint and puts the strict
policy on that response:
```java
// EpubReaderController.java
response.setHeader("Content-Security-Policy", "script-src 'none'");
```
The app shell keeps its own, more permissive policy. That works because each section is
then a real same-origin document with its own header, rather than a `blob:` — and a
`blob:` inherits the CSP of the document that created it, which is exactly why a header
on `/api/books/download/…` would achieve nothing today.
Two ways forward:
1. **Cheap.** Patch the vendored `sandbox` attribute to drop `allow-scripts`, restoring
what epub.js gave us. Cost is the WebKit bug the upstream comment cites: events inside
the iframe get swallowed, which would likely break touch/swipe paging and possibly the
in-iframe keyboard handling in `foliate-view.svelte`. Needs testing before trusting.
2. **Right.** Follow grimmory: serve individual EPUB entries from the backend with
`script-src 'none'` on each response, and drive foliate through its loader hooks
instead of a whole-file blob. This is **net-new capability on both sides**, not a
rewiring of something that exists — see below.
Option 2 also fixes the memory cost below, which is why it is worth more than it looks.
#### What option 2 actually involves
Today the browser fetches the whole `.epub` from `download/{book_id}/{file_id}`, which
returns a Litestar `File` and knows nothing about the archive's contents. foliate then
opens the zip **in the browser** (`makeZipLoader` in `view.js`) and turns every chapter,
image and stylesheet into a `blob:` URL via `Loader.createURL` in `epub.js`. A `blob:`
carries no headers of its own — it inherits the CSP of the document that created it —
which is why there is nowhere to attach a policy except the app shell.
foliate's parser never touches the zip directly. `EPUB` is constructed with a loader:
```js
// view.js — makeZipLoader is one implementation; makeDirectoryLoader below is another
return { entries, loadText, loadBlob, getSize };
```
`name` is the **zip entry path**, because `makeZipLoader` keys its map on
`entry.filename` — so `OEBPS/Text/chapter01.xhtml`, `OEBPS/Images/cover.jpg`. The parser
resolves hrefs from the OPF manifest into those names and asks the loader for them,
without caring where the bytes come from. A third implementation that fetches over HTTP
is the same shape.
**Backend.** An endpoint taking a path inside the archive, e.g.
`GET books/{book_id}/files/{file_id}/entry/{path:path}`, returning a `Stream` over
`zipfile.ZipFile.open(name)` so an entry never lands in memory whole, with the content
type from the manifest and `Content-Security-Policy: script-src 'none'` on the response.
Two things to get right:
- **`path` is caller-supplied.** Resolve it against the archive's `namelist()` and reject
anything absent, rather than trusting the string — `../` traversal is the hazard.
- **`getSize` is synchronous** in foliate's loader contract, and it feeds `SectionProgress`,
which produces the reading percentage. So the endpoint needs a companion that returns
entry names and sizes up front — one extra call at open — because sizes cannot be
discovered per request.
Opening the zip per request costs a central-directory read each time. Probably fine for
chapter-sized reads, worth measuring rather than assuming.
Note the backend already reads inside EPUBs — `metadata_extractor.py` uses ebooklib at
ingest for title, authors, identifiers and the cover. What is missing is serving an
arbitrary entry by path, not the ability to open the archive.
**Frontend.** `foliate-view.svelte` stops calling `view.open(file)` and builds an `EPUB`
around a loader backed by that endpoint. The whole-file fetch in `epub-reader.svelte`
goes away with it.
### The proxy buffers whole files and drops range headers
Two separate problems that both live in `frontend/src/routes/api/[...path]/+server.ts`.
**Buffering.** Litestar already streams: `ASGIFileResponse` reads in 1 MB chunks
(`response/file.py`), so the backend never holds a file whole. The proxy then undoes it
with `await response.arrayBuffer()`, which does not resolve until the last byte arrives —
so the whole file sits in the node process, per concurrent reader, and the browser gets
nothing until it completes. Passing `response.body` straight through restores the stream
and is a small change.
This is now **responses only**. The request side was fixed for the Calibre archive upload,
which cannot be held in memory: POST and PATCH pass `request.body` through with
`duplex: 'half'` (`bodyOf` in the same file). The response side is the same shape of fix.
**Range.** The proxy forwards only `Content-Type`, `Content-Disposition` and
`Content-Length`. It never sends the client's `Range` upstream, and would drop
`Accept-Ranges` and `Content-Range` coming back — a 206 without `Content-Range` is
broken. So range support cannot work until the proxy is fixed, whatever the backend does.
**Litestar has no range support of its own.** In 2.21.1 the only mention of 206 in the
whole package is the `HTTP_206_PARTIAL_CONTENT` constant; there is no `Accept-Ranges` or
`Content-Range` handling anywhere. This has to be written.
#### What pdf.js actually needs
It decides from the **initial 200 response**, not from anything on a 206.
`validateRangeRequestCapabilities` in `frontend/static/pdfjs/build/pdf.mjs`:
```js
if (responseHeaders.get("Accept-Ranges") !== "bytes") {
return returnValues; // allowRangeRequests stays false
}
```
It also needs a parseable `Content-Length`, `Content-Encoding: identity`, and a length
greater than twice `rangeChunkSize`. Miss any of those and it downloads the whole file
however good the range support is.
So the single highest-value header is **`Accept-Ranges: bytes` on the ordinary 200** —
that is what makes pdf.js switch to fetching progressively at all.
#### Approach
Put it on the existing `get_file` handler in `controllers/book.py`, which already resolves
`book_id`/`file_id` through the service with library scoping and auth:
- No `Range``Stream` the file with `Accept-Ranges: bytes` and `Content-Length`.
- `Range` present → parse, seek, `Stream` with 206 and `Content-Range`.
- Proxy: forward `Range` up; pass `response.body` through; forward `Accept-Ranges`,
`Content-Range` and the status back.
There are `RangeRequestMiddleware` snippets circulating for Litestar that wrap
`create_static_files_router`. They are the wrong shape here — book files are served by an
authenticated handler resolving database ids, not by a directory mapping, and using one
would mean exposing disk paths as URLs and re-solving ownership checks that already
exist. The common version also only sets `Accept-Ranges` on the 206, so it would not
switch pdf.js over, and it derives its path with `str.lstrip(prefix)`, which strips a
character set rather than a prefix — `/static/castle.pdf` becomes `le.pdf`. Worth reading
its `parse_range_header` for the parsing rules and writing the rest fresh.
### Remove the epub.js locations-cache purge
`frontend/src/lib/reader/legacy-cache.ts``purgeLegacyLocationCache`
The epub.js reader cached generated locations in `localStorage` under
`${bookId}-locations`, a few hundred KB of JSON per long book. foliate computes
progress from section byte sizes at open time, so nothing writes those keys any
more, but existing browsers still hold them — and a reader near the 510 MB
origin quota would make the new reader-settings write throw `QuotaExceededError`.
The reader clears them once per browser, behind a `chitai:locations-purged` flag.
Delete the module, its call in `epub-reader.svelte` and the flag once deployments
have had a release or two to run it — after roughly 2026-12.
### Cover dimensions are unknown until load
`book-cover.svelte` renders covers at a fixed height with natural width so nothing is
cropped or distorted. Because the intrinsic size is not known ahead of time, the text
beside the cover settles once when the image loads. `min-width` bounds the movement but
does not remove it.
Removing it properly means storing cover dimensions at ingest and emitting them as
`width`/`height` attributes so the browser reserves exact space. That is a model change
plus a migration.
### No series navigation
`Book` has `series` and `series_position`, and the detail page shows both, but there is no
way to reach the other volumes. The books list endpoint filters by author, publisher, tag,
shelf and progress — a `series` filter would need adding in
`backend/src/chitai/services/filters/book.py` and wiring through
`services/dependencies.py`, following the existing `AuthorFilter` pattern.
### Book pages are sparse for most books
EPUB metadata is thin: most imports arrive with a title, an author and nothing else. Two
independent directions, neither started:
- **Use what exists.** "More by this author" (the list endpoint already accepts
`authors=`), and exposing `Book.created_at` and `BookProgress.updated_at` — both are in
the database, neither is on `BookRead` / `BookProgressRead`.
- **Fetch from outside.** Open Library or Google Books lookup by ISBN for descriptions and
covers. This is what actually fixes the sparseness, but it needs outbound requests, rate
limiting, a manual-vs-automatic decision, and a rule for not clobbering hand-edited
metadata.
+1
View File
@@ -12,4 +12,5 @@ wheels/
# Project specific directories/files
covers/
books/
libraries/
.postgres/
+388
View File
@@ -0,0 +1,388 @@
# Chitai backend
Litestar REST API for the eBook library. See the repo-root `AGENTS.md` for the overall picture and
dev-environment setup.
**Stack:** Python 3.13 · Litestar 2 · advanced-alchemy over async SQLAlchemy 2 · asyncpg ·
PostgreSQL 17 · Alembic (through advanced-alchemy's `alchemy` CLI) · pydantic-settings · uv.
## Layering
```
controllers/ HTTP surface only — parse, delegate, serialise. Keep thin.
services/ Business logic. One SQLAlchemyAsyncRepositoryService subclass per aggregate.
database/models/ SQLAlchemy models.
schemas/ Pydantic DTOs for request bodies and responses.
```
Controllers call **service** methods, not the repository. `app.py` assembles the app: route
handlers, JWT auth, exception handlers, the SQLAlchemy plugin, and two lifespan context managers.
## advanced-alchemy idioms
These are the conventions that are easy to get wrong if you write plain SQLAlchemy here:
- Models extend `BigIntAuditBase` (adds id/created_at/updated_at); pure link and child rows extend
`BigIntBase` (e.g. `Identifier`, `FileMetadata`, `BookAuthorLink`).
- A service declares an inner repository and points at it:
```python
class BookService(SQLAlchemyAsyncRepositoryService[Book]):
class Repo(SQLAlchemyAsyncRepository[Book]):
model_type = Book
repository_type = Repo
```
- Transform incoming data with the **`to_model_on_create` / `to_model_on_update` hooks**, not by
overriding `create`/`update` wholesale — see `services/book.py:407` onward.
- Serialise responses through `service.to_schema(obj, schema_type=s.SomeRead)`; for lists,
`to_schema(items, total, filters, schema_type=…)` produces the `OffsetPagination` envelope.
- `Author`, `Tag`, `Publisher` and `BookSeries` are deduplicated with **`as_unique_async`**. Never
construct them directly when attaching to a book — use
`await Author.as_unique_async(session, name=name)` as
`BookService._populate_with_unique_relationships` does, or you will create duplicate rows.
- Domain methods on services are named for the domain (`create_book`, `update_book`, `add_files`),
deliberately distinct from the inherited CRUD names.
## Dependency injection
`services/dependencies.py` is the hub; controllers wire providers in their `dependencies` dict.
- Most providers come from `create_service_provider(SomeService, …)` — one line each.
- `provide_book_service` is hand-written because it must inject eager loads *and* scope
user-specific rows: `selectinload` for authors/tags/files/etc., plus `with_loader_criteria` so
`BookProgress` and `BookListLink` only load rows belonging to `current_user`. If you add a
relationship that the API returns, add it to that `load` list.
- `create_book_filter_dependencies` intentionally **overrides** advanced-alchemy's stock providers:
the search filter becomes a trigram search, and order-by gains a `random` sort order. Do not
replace it with the stock `create_filter_dependencies`.
- `get_library_by_id` resolves the target library from either a `library_id` query param or the
book's own `library_id`, and raises 404 for either miss.
## Filters
`services/filters/` holds `StatementFilter` dataclasses that compose into any `list` /
`list_and_count` call: `TagFilter`, `AuthorFilter`, `BookshelfFilter`, `ProgressFilter`,
`TrigramSearchFilter`, `CustomOrderBy`, `FileHashFilter`, plus the `*LibraryFilter` variants used by
OPDS.
To add list-filtering behaviour: write the dataclass here, add a `provide_*_filter` function in
`dependencies.py`, register it in the controller's `dependencies`, and fold it into
`provide_book_filters` — the controller handler itself does not change.
## Authentication — three schemes
| Scheme | Where | Used by |
| --- | --- | --- |
| JWT bearer (`OAuth2PasswordBearerAuth`) | `app.py` | The web frontend and the main API. Public paths are listed in its `exclude=[…]`. |
| HTTP Basic (`middleware/basic_auth.py`) | `OpdsController` | E-reader / OPDS clients, which only speak Basic. |
| `X-AUTH-USER` API key (`middleware/kosync_auth.py`) | `KosyncController` | KOReader devices; the key maps to a `KosyncDevice` row, which maps to a user. |
Each middleware resolves a `User` onto the connection; the matching
`provide_user_via_basic_auth` / `provide_user_via_kosync_auth` dependencies expose it to handlers.
## Filesystem behaviour
The backend owns files on disk, not just rows:
- **Layout** — `services/filesystem_library.py` (`BookPathGenerator`) renders a Jinja2 template
against book metadata to decide where a book lives under the library's `root_path`
(default: `author/series/position - title/`).
- **One directory per book, never shared.** The generated path is a pure function of the metadata,
so two books with the same author and title produce the same one — two editions, or an
`allow_duplicates` copy. `BookService._reserve_book_path` moves the later one to `title (2)`
before anything is written, and `update_book` reserves the same way so a rename cannot move a
book in on top of another. This matters because `book.path` is what deletes, moves and file
lookups act on: books sharing a directory means one overwrites the other's files, and deleting
either takes both. `_unused_path` does the same job for filenames within a directory.
A book that already has a `path` keeps it — `add_files` must follow the book, not the template.
- **Metadata extraction** — `services/metadata_extractor.py` reads EPUB (ebooklib) and PDF
(pypdfium2) files; extracted values fill only *empty* fields on the incoming payload.
- **Covers** — converted to WebP with a UUID filename under `settings.book_cover_path`, served by a
static-files router mounted at `/covers`.
- **Consume directory** — `services/consume.py` (`ConsumeDirectoryWatcher`) watches
`settings.consume_path` with `watchfiles`, creates one subdirectory per library slug, batches
additions (3 s debounce) and imports them via `BookService.create_many_from_existing_files`.
Started as an asyncio task from the `setup_directory_watcher` lifespan hook.
- **Updates move files.** `BookService.update_book` regenerates the path from the new metadata and,
if it differs, moves the directory contents and prunes empty parents. Keep that in mind before
changing metadata handling.
### Content types come from the extension, and null means null
Every ingest path names a file's format with `guess_content_type` (`services/utils.py`), never with
`mimetypes.guess_type` directly and never with what the client said. Python's built-in map answers
`None` for `.mobi`, `.azw`, `.prc`, `.fb2`, `.fbz`, `.lit`, `.lrf` and `.cb7` — most of what a
library imported from elsewhere carries — so `EBOOK_CONTENT_TYPES` fills those in. A browser's
`application/octet-stream` is discarded rather than used as a fallback: it is the client saying it
does not know, and storing it is indistinguishable from having determined a format.
When nothing can name the extension the column **stays null**. That is the honest answer, and only
one consumer cannot take it: OPDS `Link.type` is a required string, so
`services/opds/opds.py` substitutes `application/octet-stream` at that boundary. Litestar's
`ASGIFileResponse` already does its own fallback, so `get_file` can pass a null straight through.
`FileMetadataRead.content_type` is nullable for the same reason — it was once required, which
turned a stored null into a 500 on a book that was otherwise fine.
## Duplicate detection
Every ingest path screens incoming files against what is already stored, keyed on
**`(hash, size)`** — never the hash alone, because it samples 12 KiB (see below) and
EPUBs from one toolchain often share their first window. `FileMetadata.hash` carries a
plain, deliberately **non-unique** index: a collision must not be able to fail an import,
and older databases may already hold duplicates.
Scope comes from `CHITAI_DUPLICATE_SCOPE` (`library`, the default | `global` | `off`).
The policy differs by how deliberate the import is:
| Path | Behaviour |
| --- | --- |
| `create_many_from_files` (browser bulk) | Skip per file, skip a whole group whose files are all known, report everything skipped in `ImportResult.duplicates`. Re-dropping a folder to pick up what is new is the case this serves. |
| `create_book` (single, with metadata) | All-or-nothing: raises `DuplicateFilesError`, which `controllers/book.py` renders as a **409** carrying the refused files in `extra`. |
| `add_files` | A file the book already carries is a no-op; one stored under another book raises `DuplicateFilesError`. |
| `create_many_from_existing_files` (consume watcher) | Skips, and **moves the file to `CHITAI_DUPLICATE_PATH/<library slug>/`** — nothing is deleted, and it cannot stay put because `watchfiles` only reports additions. That path must stay outside `consume_path` or the watcher re-imports it and tries to read the directory name as a library slug. |
| `create_many_from_calibre` (Calibre import) | Skips per file, and skips a whole book whose files are all known. Nothing is moved — the source is somebody else's library. This is what makes a re-run a no-op and an interrupted import resumable by running it again. |
`allow_duplicates=true` overrides all of it, on every endpoint. Keep that working — the
hash is not proof of identity, so a wrong verdict has to be recoverable, and a scripted
import needs a way through. The **web UI deliberately does not offer it**: storing the
same bytes twice splits reading progress and shelf membership across two records that
can never converge, which is nothing anyone wants on purpose.
Three things to preserve when touching this code:
- **A match only counts while the file is on disk.** `find_duplicate_files` stats each
candidate, and `add_files` writes a missing file back into the row that already
describes it (`_restore_file`) instead of adding a second row beside it. The hash
lives in the database and the file does not, so without this a file deleted behind
the app's back would go on refusing its own replacement.
- **Screening runs before anything is written.** `fingerprint_upload` reads the spooled
upload and rewinds it; the resulting fingerprints are handed to `_save_book_files`,
which skips its own `StreamingHasher` when it already has the answer. Passing them
through is what keeps the file from being read twice.
- **`_screen_for_duplicates` extends the `known` dict as it goes**, so the same bytes
submitted twice in one request are caught. Those duplicates report `book_id: None` —
there is no row to point at yet.
`POST /books/duplicate-files` answers the same question from fingerprints alone, for
clients that want to ask before uploading anything.
### Book-level detection — a different question
The file check answers "are these the same bytes?". `find_duplicate_books` answers "is
this the same book?", which a re-scan, a re-zipped EPUB or another edition cannot be
asked with a hash. Two signals, either sufficient: a **shared identifier**, or a
**matching normalized title with at least one shared author**.
It **never blocks**. A metadata match is a guess — a work shares title and author with
its own translation, its own second edition and its own audiobook — so the book is
created and the candidates are reported alongside it in
`ImportResult.possible_duplicates`. File-level dedupe keeps its refuse/skip behaviour;
that one is near-certain and this one is not. Do not "improve" this into a refusal.
Two rules narrow it, both applied in Python over the small candidate set:
- **A shared author is required for a title match.** Without it every book the
extractors gave up on and titled `Unknown` is a duplicate of every other one. A book
with no authors can therefore only match on an identifier.
- **The same series at a different `series_position` disqualifies a match.** A trilogy
shares an author and often most of its title; the position is the library saying
outright that these are two books.
A book is compared under **several title keys, not one** (`_title_keys`). Ebook files
are overwhelmingly named `Title - Author.epub`, and wherever nothing inside the file
overrode that name the author ended up in the title column — so one copy is stored as
`Building Microservices` and another as `Building Microservices Sam Newman`. Both
directions are generated, the author stripped off and the author added on, which is why
the query is `normalized_title.in_(keys)` rather than `==`. This is still exact matching
on an indexed column: no similarity score, nothing to tune. It does not weaken the
shared-author requirement, which is a separate condition.
`find_duplicate_book_groups` is the library-wide pass behind
`GET /books/duplicate-books`, since the import-time check says nothing about a
collection someone already has. It buckets books by every key they carry and merges the
buckets with union-find, so A~B by ISBN and B~C by title land in one group. Pairs in
`duplicate_dismissals` are never merged — a reader disagreeing with one pairing must not
silently break a group that stands on other evidence.
### Author names have one stored form
`Author.name` is always the canonical form, produced by `format_author_name`. Extractors
hand over whatever the file said — `Newman, Sam;` from a `DC:creator` list, `Sam Newman`
from a PDF, `Sam Newman.epub` from a filename — and storing those verbatim is how one
person becomes four rows in the sidebar, four entries in the author filter, and four
books that never look like each other.
This is **display** canonicalization, distinct from `normalize_author`, which throws
away case, accents and spacing to build a comparison key nobody sees. Tidying only
removes what an extractor added: a trailing separator, a file extension, and the
`Surname, Given` ordering. It never touches case or accents — `Michał Płachta` and
`Steve McConnell` are the author's own spelling, not something to correct.
Three places have to agree, and `Author` keeps them together: the `@validates("name")`
hook, `unique_hash`, and `unique_filter`. `as_unique_async` looks a row up with the
filter and then constructs with the validator, so if the lookup used the raw name and
the insert used the tidy one, every variant spelling would miss the existing row and
then collide with it on the unique index.
A form the rule does not recognise is **left exactly as it was found** — `Dave Thomas,
Andy Hunt` is two people in one string, and flipping it would invent a third. Leaving a
mess visible beats rewriting it wrongly.
### The normalized columns are written by validators
`Book.normalized_title`, `Author.normalized_name` and `Identifier.normalized_value` are
derived from `services/matching.py` and kept current by **SQLAlchemy `@validates` hooks
on the models**, not by any service. Assigning them directly is always wrong.
This is deliberate and it is invisible at the call sites: `BookService` writes titles
through at least three paths (`to_model_on_create`, `to_model_on_update`, and the
`setattr` loop in `_populate_with_unique_relationships`), and a validator is the only
thing a fourth cannot bypass. The columns carry plain, **non-unique** btree indexes —
two spellings collapsing onto one value is the entire point.
`services/matching.py` is imported *inside* those validators rather than at module
scope: reaching it initialises the `chitai.services` package, which imports the
services, which import the models. Keep the local import.
The validators only fire on write, so a migration that adds one of these columns must
backfill existing rows through the same helpers — see the `data_upgrades()` hook in
`2026-08-15_add_book_matching_keys_and_duplicate__4358e7d4743a.py`.
**Changing anything in `services/matching.py` needs a revision that recomputes them.**
The keys are derived and already written, so a normalization change silently invalidates
every stored row: a book written under the old rules just stops matching one written
under the new rules, with nothing to show that anything is wrong. Copy
`2026-08-15_recompute_book_matching_keys_ed41acf21270.py`, which exists because
`normalize_title` learned to strip compact edition markers (`2E`, `5e`). It is
idempotent and safe to re-run.
## Importing from Calibre
Two pieces, deliberately separated:
- **`services/calibre.py`** reads `metadata.db` and the tree beside it. It knows nothing about
`Book`, `BookService` or a session, so it is testable without Postgres, and it reports what
Calibre wrote rather than what Chitai wants — identifiers come back keyed by `identifiers.type`
verbatim. It also unpacks a zipped library (`extract_calibre_archive`).
- **`BookService.create_many_from_calibre`** does the ingest, next to the two other ingest paths
because it needs the same privates they do (`_reserve_book_path`, `_save_cover_image`,
`_screen_for_duplicates`). The CLI in `cli.py` is a thin wrapper over it.
- **`services/calibre_import.py`** is lifecycle only — the job registry behind the endpoints:
state, progress, cancellation, and a session of its own.
`docs/calibre-import.md` is the full brief, including the phases not built yet. What matters here:
- **Files are copied, never moved.** `metadata.db` would go on pointing at files that are gone,
which quietly ruins a library somebody still uses. `copy_file` streams rather than using
`shutil.copy`, which would block the loop for a 40 MB read.
- **The extractors are not run.** This is the one ingest path that trusts its input: Calibre's
catalogue is curated and its filenames are truncated to ~42 characters, so `data.name` locates a
file and the database carries the metadata. The title is stored verbatim for the same reason —
no edition split out, no subtitle guessed.
- **The catalogue is copied before it is read**, and the copy is opened read-write. Calibre may be
running; opening the live file either sees a torn state or needs to recover a write-ahead log,
which read-only access cannot do. `close()` removes the copy in a `finally`, or a failure leaves
a catalogue-sized file in the temp directory.
- **`check_same_thread=False` plus an `asyncio.Lock`.** Every query runs through
`asyncio.to_thread`, which hands out whichever worker is free, so the connection outlives the
thread that opened it. The lock is what makes that safe. Removing either one reintroduces
`SQLite objects created in a thread can only be used in that same thread`, intermittently —
the pool often reuses one thread, so it passes until it does not.
- **One book never costs the run.** A failure is recorded in `CalibreImportResult.failed`, the
session is rolled back so the next book can use it, and the files that book had already copied
are deleted — an orphaned directory would make the next attempt reserve `title (2)` and look as
though it had worked. A cover that PIL cannot open costs the cover, not the book.
- **`calibre-uuid`, not `uuid`.** `books.uuid` is stable for the life of the row, so it is the
durable link back to the source and worth matching on. `uuid` is the name `normalize_identifier`
refuses, because an EPUB regenerates one per build.
### The API takes an uploaded archive; the CLI takes a path
**`POST /libraries/{id}/imports/calibre/upload`** is the only way in over HTTP. It takes a zipped
Calibre library, answers **202** with a job handle, and unpacks into a temp directory the job owns.
`GET /libraries/imports/{job_id}` is polled; `DELETE` on the same path stops it.
There is deliberately **no endpoint that imports from a server path**. A desktop Calibre install is
not on the server, and importing from a path the server can already see is a server-side operation
— which is what `litestar --app-dir src/chitai/ calibre-import <path> --library <slug>` is for,
including its `--dry-run`. Do not add the path endpoint back without being asked: it was built,
then removed on purpose.
- **The registry is in memory, so it assumes one worker process.** That holds today (`litestar run`
is single-process, and the consume watcher is already an in-process singleton), but the day
`TODO.md`'s "production image runs the development server" item is fixed with a worker count, a
poll can land on a worker that never heard of the job. `services/calibre_import.py` says so at
the top; an `import_jobs` table is the answer when that happens.
- **The job opens its own session.** The request that started it is long gone and its session
closed with it.
- **Cancelling is not aborting.** A flag is read between books, never during one, so a cancelled
import leaves whole books behind and never half of one. `task.cancel()` would abandon a book
mid-copy and leave files with no row describing them.
- **The job deletes its workspace** — the unpacked archive is a second copy of the whole library,
and the books worth keeping have been copied into the library proper by the time it ends. Removed
even when the run failed, since nothing will come back for it.
- **Extraction refuses** an entry pointing outside the archive (zip slip), an archive that will not
fit on disk, and one with no `metadata.db` within three levels. All three answer 400 before a job
exists, rather than as a job that reports FAILED a moment later.
- **The upload is streamed both sides.** The archive reaches disk in chunks rather than being read
whole, and the SvelteKit proxy passes `request.body` through instead of buffering it — see
`frontend/AGENTS.md`.
Things Calibre does that will produce wrong data if you forget them are documented at the top of
`services/calibre.py` — the `0101-01-01` date sentinel, `|` for a comma in an author name, the
REAL `series_index` that defaults to 1.0 for every book, HTML in `comments`, the views that need
SQLite functions Calibre registers from Python, and `books_pages_link` being both recent and
usually empty. `tests/calibre_fixtures.py` builds a library exercising all of them.
## KOReader hashing
`services/utils.py` reimplements KOReader's partial-MD5 document identifier: 1 KiB samples at
offsets produced by LuaJIT's 32-bit `bit.lshift`, including its shift-masking wrap-around
(`shift & 0x1F`, so `i=-1` yields offset 0). `_lshift32` looks wrong and is not — the overflow is
what makes hashes match real devices. Don't "simplify" it; `tests/integration/test_file_hash.py`
guards the behaviour.
## Database
- **`pg_trgm` is required.** `Book.__table_args__` declares a GIN trigram index on `title`, which
`TrigramSearchFilter` uses for fuzzy title search. The extension is enabled by a migration and, in
tests, by `conftest.py`.
- Migrations live in `migrations/versions/`. Generate with
`alchemy --config chitai.database.config.config make-migrations` (add `--no-autogenerate` for a
blank revision), apply with `… upgrade`. `database/config.py` sets `create_all=False`, so nothing
is auto-created at runtime; production applies migrations from `entrypoint.sh`.
- Sessions use `expire_on_commit=False` and Litestar's `before_send_handler="autocommit"`, so a
handler that returns 2xx commits automatically.
## Testing
`pytest tests/` — `asyncio_mode = "auto"`, so async tests need no marker.
- `tests/unit/` exercises services directly against a real session; `tests/integration/` drives the
whole app through `AsyncTestClient(app=create_app())`.
- The database is a throwaway container from `pytest-databases`, so **Docker must be running**.
- `tests/conftest.py` provides the shared fixtures: `client`, `authenticated_client`,
`other_authenticated_client` (a second user, for access-control tests), one fixture per service,
`test_user` / `test_library`, and an autouse fixture that redirects cover storage into `tmp_path`.
- `tests/integration/conftest.py` monkeypatches the module-level alchemy `config` onto the test
engine/sessionmaker and drops+recreates+reseeds the schema for every test.
- Real EPUB and PDF fixtures live in `tests/data_files/`.
## Known rough edges
Observed in the current tree — don't mistake these for intentional patterns to copy:
- `controllers/book.py` — file-level TODO: `book_id` is a path parameter on some endpoints and a
query parameter on others. `set_book_progress_batch` does a documented N+1 (one select + one
upsert per book).
- `services/filesystem_library.py` — TODO to replace Jinja2 templating with simple placeholders;
`generate_filename` accepts a `filename_template` but currently ignores it and returns the
original filename.
- `app.py` — `watcher_task` is declared as a module-level global but assigned locally inside
`setup_directory_watcher`, so the global is never populated (cancellation still works via the
closure).
- `BookService.get_files` (the multi-book ZIP download) opens `Path(file.path)`, but `file.path` is
stored relative to `book.path` — worth verifying before relying on that endpoint.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+1
View File
@@ -0,0 +1 @@
Asynchronous SQLAlchemy configuration with Advanced Alchemy.
+1 -1
View File
@@ -1,12 +1,12 @@
import asyncio
from typing import TYPE_CHECKING, cast
from alembic.autogenerate import rewriter
from sqlalchemy import pool
from sqlalchemy.ext.asyncio import AsyncEngine, async_engine_from_config
from advanced_alchemy.base import metadata_registry
from alembic import context
from alembic.autogenerate import rewriter
if TYPE_CHECKING:
from sqlalchemy.engine import Connection
@@ -1,8 +1,8 @@
"""Add pg_trgm extension
"""add postgres extensions
Revision ID: 26022ec86f32
Revision ID: 65aa95e8f8cf
Revises:
Create Date: 2025-10-31 18:45:55.027462
Create Date: 2026-03-14 11:59:55.364307
"""
@@ -11,7 +11,12 @@ from typing import TYPE_CHECKING
import sqlalchemy as sa
from alembic import op
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash, FernetBackend
from advanced_alchemy.types.encrypted_string import PGCryptoBackend
from advanced_alchemy.types.password_hash.argon2 import Argon2Hasher
from advanced_alchemy.types.password_hash.passlib import PasslibHasher
from advanced_alchemy.types.password_hash.pwdlib import PwdlibHasher
from pwdlib.hashers.argon2 import Argon2Hasher as PwdlibArgon2Hasher
from sqlalchemy import Text # noqa: F401
if TYPE_CHECKING:
@@ -26,17 +31,22 @@ sa.EncryptedString = EncryptedString
sa.EncryptedText = EncryptedText
sa.StoredObject = StoredObject
sa.PasswordHash = PasswordHash
sa.Argon2Hasher = Argon2Hasher
sa.PasslibHasher = PasslibHasher
sa.PwdlibHasher = PwdlibHasher
sa.FernetBackend = FernetBackend
sa.PGCryptoBackend = PGCryptoBackend
# revision identifiers, used by Alembic.
revision = '26022ec86f32'
revision = '65aa95e8f8cf'
down_revision = None
branch_labels = None
depends_on = None
def upgrade() -> None:
op.execute(sa.text('create EXTENSION if not EXISTS "pgcrypto"'))
op.execute(sa.text('create EXTENSION if not EXISTS "pg_trgm"'))
op.execute(sa.text('CREATE EXTENSION IF NOT EXISTS "pgcrypto"'))
op.execute(sa.text('CREATE EXTENSION IF NOT EXISTS "pg_trgm"'))
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
@@ -0,0 +1,298 @@
"""create initial tables
Revision ID: 6d72d1bbc0ee
Revises: 65aa95e8f8cf
Create Date: 2026-03-14 14:36:43.529584
"""
import warnings
from typing import TYPE_CHECKING
import sqlalchemy as sa
from alembic import op
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash, FernetBackend
from advanced_alchemy.types.encrypted_string import PGCryptoBackend
from advanced_alchemy.types.password_hash.argon2 import Argon2Hasher
from advanced_alchemy.types.password_hash.passlib import PasslibHasher
from advanced_alchemy.types.password_hash.pwdlib import PwdlibHasher
from pwdlib.hashers.argon2 import Argon2Hasher as PwdlibArgon2Hasher
from sqlalchemy import Text # noqa: F401
if TYPE_CHECKING:
from collections.abc import Sequence
__all__ = ["downgrade", "upgrade", "schema_upgrades", "schema_downgrades", "data_upgrades", "data_downgrades"]
sa.GUID = GUID
sa.DateTimeUTC = DateTimeUTC
sa.ORA_JSONB = ORA_JSONB
sa.EncryptedString = EncryptedString
sa.EncryptedText = EncryptedText
sa.StoredObject = StoredObject
sa.PasswordHash = PasswordHash
sa.Argon2Hasher = Argon2Hasher
sa.PasslibHasher = PasslibHasher
sa.PwdlibHasher = PwdlibHasher
sa.FernetBackend = FernetBackend
sa.PGCryptoBackend = PGCryptoBackend
# revision identifiers, used by Alembic.
revision = '6d72d1bbc0ee'
down_revision = '65aa95e8f8cf'
branch_labels = None
depends_on = None
def upgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
schema_upgrades()
data_upgrades()
def downgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
data_downgrades()
schema_downgrades()
def schema_upgrades() -> None:
"""schema upgrade migrations go here."""
# ### commands auto generated by Alembic - please adjust! ###
op.create_table('authors',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('name', sa.String(), nullable=False),
sa.Column('description', sa.String(), nullable=True),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.PrimaryKeyConstraint('id', name=op.f('pk_authors'))
)
with op.batch_alter_table('authors', schema=None) as batch_op:
batch_op.create_index(batch_op.f('ix_authors_name'), ['name'], unique=True)
op.create_table('book_series',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('title', sa.String(), nullable=False),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.PrimaryKeyConstraint('id', name=op.f('pk_book_series'))
)
with op.batch_alter_table('book_series', schema=None) as batch_op:
batch_op.create_index(batch_op.f('ix_book_series_title'), ['title'], unique=True)
op.create_table('libraries',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('name', sa.String(), nullable=False),
sa.Column('root_path', sa.String(), nullable=False),
sa.Column('path_template', sa.String(), nullable=False),
sa.Column('description', sa.String(), nullable=True),
sa.Column('icon', sa.String(), nullable=False),
sa.Column('read_only', sa.Boolean(), nullable=False),
sa.Column('slug', sa.String(length=100), nullable=False),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.PrimaryKeyConstraint('id', name=op.f('pk_libraries')),
sa.UniqueConstraint('name', name=op.f('uq_libraries_name')),
sa.UniqueConstraint('slug', name='uq_libraries_slug')
)
with op.batch_alter_table('libraries', schema=None) as batch_op:
batch_op.create_index('ix_libraries_slug_unique', ['slug'], unique=True)
op.create_table('publishers',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('name', sa.String(), nullable=False),
sa.Column('description', sa.String(), nullable=True),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.PrimaryKeyConstraint('id', name=op.f('pk_publishers'))
)
op.create_table('tags',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('name', sa.String(), nullable=False),
sa.PrimaryKeyConstraint('id', name=op.f('pk_tags'))
)
with op.batch_alter_table('tags', schema=None) as batch_op:
batch_op.create_index(batch_op.f('ix_tags_name'), ['name'], unique=True)
op.create_table('users',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('email', sa.String(), nullable=False),
sa.Column('password', sa.PasswordHash(backend=sa.PwdlibHasher(PwdlibArgon2Hasher()), length=128), nullable=False),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.PrimaryKeyConstraint('id', name=op.f('pk_users')),
sa.UniqueConstraint('email', name=op.f('uq_users_email'))
)
op.create_table('book_lists',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('library_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=True),
sa.Column('user_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('title', sa.String(), nullable=False),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.ForeignKeyConstraint(['library_id'], ['libraries.id'], name=op.f('fk_book_lists_library_id_libraries')),
sa.ForeignKeyConstraint(['user_id'], ['users.id'], name=op.f('fk_book_lists_user_id_users')),
sa.PrimaryKeyConstraint('id', name=op.f('pk_book_lists'))
)
op.create_table('books',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('library_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('title', sa.String(), nullable=False),
sa.Column('subtitle', sa.String(), nullable=True),
sa.Column('description', sa.String(), nullable=True),
sa.Column('published_date', sa.Date(), nullable=True),
sa.Column('language', sa.String(), nullable=True),
sa.Column('pages', sa.Integer(), nullable=True),
sa.Column('cover_image', sa.String(), nullable=True),
sa.Column('edition', sa.Integer(), nullable=True),
sa.Column('path', sa.String(), nullable=True),
sa.Column('publisher_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=True),
sa.Column('series_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=True),
sa.Column('series_position', sa.String(), nullable=True),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.ForeignKeyConstraint(['library_id'], ['libraries.id'], name=op.f('fk_books_library_id_libraries')),
sa.ForeignKeyConstraint(['publisher_id'], ['publishers.id'], name=op.f('fk_books_publisher_id_publishers')),
sa.ForeignKeyConstraint(['series_id'], ['book_series.id'], name=op.f('fk_books_series_id_book_series')),
sa.PrimaryKeyConstraint('id', name=op.f('pk_books'))
)
with op.batch_alter_table('books', schema=None) as batch_op:
batch_op.create_index('ix_books_title_trigram', ['title'], unique=False, postgresql_using='gin', postgresql_ops={'title': 'gin_trgm_ops'})
op.create_table('devices',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('user_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('api_key', sa.String(), nullable=False),
sa.Column('name', sa.String(), nullable=False),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.ForeignKeyConstraint(['user_id'], ['users.id'], name=op.f('fk_devices_user_id_users')),
sa.PrimaryKeyConstraint('id', name=op.f('pk_devices')),
sa.UniqueConstraint('api_key', name=op.f('uq_devices_api_key'))
)
op.create_table('book_author_links',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('author_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('position', sa.Integer(), nullable=False),
sa.ForeignKeyConstraint(['author_id'], ['authors.id'], name=op.f('fk_book_author_links_author_id_authors')),
sa.ForeignKeyConstraint(['book_id'], ['books.id'], name=op.f('fk_book_author_links_book_id_books'), ondelete='cascade'),
sa.PrimaryKeyConstraint('id', name=op.f('pk_book_author_links')),
sa.UniqueConstraint('book_id', 'author_id', name=op.f('uq_book_author_links_book_id'))
)
op.create_table('book_list_links',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('list_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('position', sa.Integer(), nullable=False),
sa.ForeignKeyConstraint(['book_id'], ['books.id'], name=op.f('fk_book_list_links_book_id_books'), ondelete='cascade'),
sa.ForeignKeyConstraint(['list_id'], ['book_lists.id'], name=op.f('fk_book_list_links_list_id_book_lists')),
sa.PrimaryKeyConstraint('id', name=op.f('pk_book_list_links')),
sa.UniqueConstraint('book_id', 'list_id', name=op.f('uq_book_list_links_book_id'))
)
op.create_table('book_progress',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('user_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('epub_cfi', sa.String(), nullable=True),
sa.Column('epub_xpointer', sa.String(), nullable=True),
sa.Column('pdf_page', sa.Integer(), nullable=True),
sa.Column('percentage', sa.Float(), nullable=False),
sa.Column('completed', sa.Boolean(), nullable=True),
sa.Column('device', sa.String(), nullable=True),
sa.Column('device_id', sa.String(), nullable=True),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.ForeignKeyConstraint(['book_id'], ['books.id'], name=op.f('fk_book_progress_book_id_books'), ondelete='cascade'),
sa.ForeignKeyConstraint(['user_id'], ['users.id'], name=op.f('fk_book_progress_user_id_users'), ondelete='cascade'),
sa.PrimaryKeyConstraint('id', name=op.f('pk_book_progress'))
)
op.create_table('book_tag_link',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('tag_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('position', sa.Integer(), nullable=False),
sa.ForeignKeyConstraint(['book_id'], ['books.id'], name=op.f('fk_book_tag_link_book_id_books'), ondelete='cascade'),
sa.ForeignKeyConstraint(['tag_id'], ['tags.id'], name=op.f('fk_book_tag_link_tag_id_tags')),
sa.PrimaryKeyConstraint('id', name=op.f('pk_book_tag_link')),
sa.UniqueConstraint('book_id', 'tag_id', name=op.f('uq_book_tag_link_book_id'))
)
op.create_table('file_metadata',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('hash', sa.String(), nullable=False),
sa.Column('path', sa.String(), nullable=False),
sa.Column('size', sa.Integer(), nullable=False),
sa.Column('content_type', sa.String(), nullable=True),
sa.ForeignKeyConstraint(['book_id'], ['books.id'], name=op.f('fk_file_metadata_book_id_books'), ondelete='cascade'),
sa.PrimaryKeyConstraint('id', name=op.f('pk_file_metadata'))
)
op.create_table('identifiers',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('name', sa.String(), nullable=False),
sa.Column('book_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('value', sa.String(), nullable=False),
sa.ForeignKeyConstraint(['book_id'], ['books.id'], name=op.f('fk_identifiers_book_id_books'), ondelete='cascade'),
sa.PrimaryKeyConstraint('id', name=op.f('pk_identifiers')),
sa.UniqueConstraint('name', 'book_id', name=op.f('uq_identifiers_name'))
)
op.create_table('kosync_progress',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('user_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('document', sa.String(), nullable=False),
sa.Column('progress', sa.String(), nullable=True),
sa.Column('percentage', sa.Float(), nullable=True),
sa.Column('device', sa.String(), nullable=True),
sa.Column('device_id', sa.String(), nullable=True),
sa.Column('created_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.Column('updated_at', sa.DateTimeUTC(timezone=True), nullable=False),
sa.ForeignKeyConstraint(['book_id'], ['books.id'], name=op.f('fk_kosync_progress_book_id_books'), ondelete='cascade'),
sa.ForeignKeyConstraint(['user_id'], ['users.id'], name=op.f('fk_kosync_progress_user_id_users'), ondelete='cascade'),
sa.PrimaryKeyConstraint('id', name=op.f('pk_kosync_progress'))
)
# ### end Alembic commands ###
def schema_downgrades() -> None:
"""schema downgrade migrations go here."""
# ### commands auto generated by Alembic - please adjust! ###
op.drop_table('kosync_progress')
op.drop_table('identifiers')
op.drop_table('file_metadata')
op.drop_table('book_tag_link')
op.drop_table('book_progress')
op.drop_table('book_list_links')
op.drop_table('book_author_links')
op.drop_table('devices')
with op.batch_alter_table('books', schema=None) as batch_op:
batch_op.drop_index('ix_books_title_trigram', postgresql_using='gin', postgresql_ops={'title': 'gin_trgm_ops'})
op.drop_table('books')
op.drop_table('book_lists')
op.drop_table('users')
with op.batch_alter_table('tags', schema=None) as batch_op:
batch_op.drop_index(batch_op.f('ix_tags_name'))
op.drop_table('tags')
op.drop_table('publishers')
with op.batch_alter_table('libraries', schema=None) as batch_op:
batch_op.drop_index('ix_libraries_slug_unique')
op.drop_table('libraries')
with op.batch_alter_table('book_series', schema=None) as batch_op:
batch_op.drop_index(batch_op.f('ix_book_series_title'))
op.drop_table('book_series')
with op.batch_alter_table('authors', schema=None) as batch_op:
batch_op.drop_index(batch_op.f('ix_authors_name'))
op.drop_table('authors')
# ### end Alembic commands ###
def data_upgrades() -> None:
"""Add any optional data upgrade migrations here!"""
def data_downgrades() -> None:
"""Add any optional data downgrade migrations here!"""
@@ -0,0 +1,80 @@
"""add file hash index
Revision ID: e9c2c7e875ae
Revises: 6d72d1bbc0ee
Create Date: 2026-08-13 14:52:09.341906
"""
import warnings
from typing import TYPE_CHECKING
import sqlalchemy as sa
from alembic import op
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash, FernetBackend
from advanced_alchemy.types.encrypted_string import PGCryptoBackend
from advanced_alchemy.types.password_hash.argon2 import Argon2Hasher
from advanced_alchemy.types.password_hash.passlib import PasslibHasher
from advanced_alchemy.types.password_hash.pwdlib import PwdlibHasher
from sqlalchemy import Text # noqa: F401
if TYPE_CHECKING:
from collections.abc import Sequence
__all__ = ["downgrade", "upgrade", "schema_upgrades", "schema_downgrades", "data_upgrades", "data_downgrades"]
sa.GUID = GUID
sa.DateTimeUTC = DateTimeUTC
sa.ORA_JSONB = ORA_JSONB
sa.EncryptedString = EncryptedString
sa.EncryptedText = EncryptedText
sa.StoredObject = StoredObject
sa.PasswordHash = PasswordHash
sa.Argon2Hasher = Argon2Hasher
sa.PasslibHasher = PasslibHasher
sa.PwdlibHasher = PwdlibHasher
sa.FernetBackend = FernetBackend
sa.PGCryptoBackend = PGCryptoBackend
# revision identifiers, used by Alembic.
revision = 'e9c2c7e875ae'
down_revision = '6d72d1bbc0ee'
branch_labels = None
depends_on = None
def upgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
schema_upgrades()
data_upgrades()
def downgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
data_downgrades()
schema_downgrades()
def schema_upgrades() -> None:
"""schema upgrade migrations go here."""
# ### commands auto generated by Alembic - please adjust! ###
with op.batch_alter_table('file_metadata', schema=None) as batch_op:
batch_op.create_index('ix_file_metadata_hash', ['hash'], unique=False)
# ### end Alembic commands ###
def schema_downgrades() -> None:
"""schema downgrade migrations go here."""
# ### commands auto generated by Alembic - please adjust! ###
with op.batch_alter_table('file_metadata', schema=None) as batch_op:
batch_op.drop_index('ix_file_metadata_hash')
# ### end Alembic commands ###
def data_upgrades() -> None:
"""Add any optional data upgrade migrations here!"""
def data_downgrades() -> None:
"""Add any optional data downgrade migrations here!"""
@@ -0,0 +1,163 @@
"""add book matching keys and duplicate dismissals
Revision ID: 4358e7d4743a
Revises: e9c2c7e875ae
Create Date: 2026-08-15 15:09:02.708914
"""
import warnings
import sqlalchemy as sa
from alembic import op
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash, FernetBackend
from advanced_alchemy.types.encrypted_string import PGCryptoBackend
from advanced_alchemy.types.password_hash.argon2 import Argon2Hasher
from advanced_alchemy.types.password_hash.passlib import PasslibHasher
from advanced_alchemy.types.password_hash.pwdlib import PwdlibHasher
from sqlalchemy import Text # noqa: F401
__all__ = ["downgrade", "upgrade", "schema_upgrades", "schema_downgrades", "data_upgrades", "data_downgrades"]
sa.GUID = GUID
sa.DateTimeUTC = DateTimeUTC
sa.ORA_JSONB = ORA_JSONB
sa.EncryptedString = EncryptedString
sa.EncryptedText = EncryptedText
sa.StoredObject = StoredObject
sa.PasswordHash = PasswordHash
sa.Argon2Hasher = Argon2Hasher
sa.PasslibHasher = PasslibHasher
sa.PwdlibHasher = PwdlibHasher
sa.FernetBackend = FernetBackend
sa.PGCryptoBackend = PGCryptoBackend
# revision identifiers, used by Alembic.
revision = '4358e7d4743a'
down_revision = 'e9c2c7e875ae'
branch_labels = None
depends_on = None
def upgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
schema_upgrades()
data_upgrades()
def downgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
data_downgrades()
schema_downgrades()
def schema_upgrades() -> None:
"""schema upgrade migrations go here."""
# ### commands auto generated by Alembic - please adjust! ###
op.create_table('duplicate_dismissals',
sa.Column('id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_a_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.Column('book_b_id', sa.BigInteger().with_variant(sa.Integer(), 'sqlite'), nullable=False),
sa.ForeignKeyConstraint(['book_a_id'], ['books.id'], name=op.f('fk_duplicate_dismissals_book_a_id_books'), ondelete='cascade'),
sa.ForeignKeyConstraint(['book_b_id'], ['books.id'], name=op.f('fk_duplicate_dismissals_book_b_id_books'), ondelete='cascade'),
sa.PrimaryKeyConstraint('id', name=op.f('pk_duplicate_dismissals')),
sa.UniqueConstraint('book_a_id', 'book_b_id', name=op.f('uq_duplicate_dismissals_book_a_id'))
)
with op.batch_alter_table('duplicate_dismissals', schema=None) as batch_op:
batch_op.create_index(batch_op.f('ix_duplicate_dismissals_book_a_id'), ['book_a_id'], unique=False)
batch_op.create_index(batch_op.f('ix_duplicate_dismissals_book_b_id'), ['book_b_id'], unique=False)
# `server_default` so the column can be added to a table that already has rows;
# `data_upgrades` fills in the real keys immediately afterwards.
with op.batch_alter_table('authors', schema=None) as batch_op:
batch_op.add_column(sa.Column('normalized_name', sa.String(), nullable=False, server_default=''))
batch_op.create_index(batch_op.f('ix_authors_normalized_name'), ['normalized_name'], unique=False)
with op.batch_alter_table('books', schema=None) as batch_op:
batch_op.add_column(sa.Column('normalized_title', sa.String(), nullable=False, server_default=''))
batch_op.create_index(batch_op.f('ix_books_normalized_title'), ['normalized_title'], unique=False)
with op.batch_alter_table('identifiers', schema=None) as batch_op:
batch_op.add_column(sa.Column('normalized_value', sa.String(), nullable=True))
batch_op.create_index(batch_op.f('ix_identifiers_normalized_value'), ['normalized_value'], unique=False)
# ### end Alembic commands ###
def schema_downgrades() -> None:
"""schema downgrade migrations go here."""
# ### commands auto generated by Alembic - please adjust! ###
with op.batch_alter_table('identifiers', schema=None) as batch_op:
batch_op.drop_index(batch_op.f('ix_identifiers_normalized_value'))
batch_op.drop_column('normalized_value')
with op.batch_alter_table('books', schema=None) as batch_op:
batch_op.drop_index(batch_op.f('ix_books_normalized_title'))
batch_op.drop_column('normalized_title')
with op.batch_alter_table('authors', schema=None) as batch_op:
batch_op.drop_index(batch_op.f('ix_authors_normalized_name'))
batch_op.drop_column('normalized_name')
with op.batch_alter_table('duplicate_dismissals', schema=None) as batch_op:
batch_op.drop_index(batch_op.f('ix_duplicate_dismissals_book_b_id'))
batch_op.drop_index(batch_op.f('ix_duplicate_dismissals_book_a_id'))
op.drop_table('duplicate_dismissals')
# ### end Alembic commands ###
def data_upgrades() -> None:
"""
Fill the matching keys in for rows that already exist.
The validators on the models only fire when something is written, so without this
every book imported before today is invisible to duplicate detection. Run through
the same helpers the validators use, so a backfilled row and a freshly written one
are guaranteed to agree.
"""
from chitai.services.matching import (
normalize_author,
normalize_identifier,
normalize_title,
)
connection = op.get_bind()
books = connection.execute(sa.text("SELECT id, title FROM books")).fetchall()
_apply(
connection,
"UPDATE books SET normalized_title = :key WHERE id = :id",
[{"id": id, "key": normalize_title(title)} for id, title in books],
)
authors = connection.execute(sa.text("SELECT id, name FROM authors")).fetchall()
_apply(
connection,
"UPDATE authors SET normalized_name = :key WHERE id = :id",
[{"id": id, "key": normalize_author(name)} for id, name in authors],
)
identifiers = connection.execute(
sa.text("SELECT id, name, value FROM identifiers")
).fetchall()
_apply(
connection,
"UPDATE identifiers SET normalized_value = :key WHERE id = :id",
[
{"id": id, "key": normalize_identifier(name, value)}
for id, name, value in identifiers
],
)
def _apply(connection, statement: str, parameters: list[dict]) -> None:
"""Run one update per row, in batches, skipping the work when there are none."""
batch_size = 1000
for start in range(0, len(parameters), batch_size):
connection.execute(sa.text(statement), parameters[start : start + batch_size])
def data_downgrades() -> None:
"""Add any optional data downgrade migrations here!"""
@@ -0,0 +1,130 @@
"""canonicalize author names
Revision ID: 49a9e85a0ffc
Revises: ed41acf21270
Create Date: 2026-08-15 15:59:47.331545
"""
import warnings
import sqlalchemy as sa
from alembic import op
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash, FernetBackend
from advanced_alchemy.types.encrypted_string import PGCryptoBackend
from advanced_alchemy.types.password_hash.argon2 import Argon2Hasher
from advanced_alchemy.types.password_hash.passlib import PasslibHasher
from advanced_alchemy.types.password_hash.pwdlib import PwdlibHasher
from sqlalchemy import Text # noqa: F401
__all__ = ["downgrade", "upgrade", "schema_upgrades", "schema_downgrades", "data_upgrades", "data_downgrades"]
sa.GUID = GUID
sa.DateTimeUTC = DateTimeUTC
sa.ORA_JSONB = ORA_JSONB
sa.EncryptedString = EncryptedString
sa.EncryptedText = EncryptedText
sa.StoredObject = StoredObject
sa.PasswordHash = PasswordHash
sa.Argon2Hasher = Argon2Hasher
sa.PasslibHasher = PasslibHasher
sa.PwdlibHasher = PwdlibHasher
sa.FernetBackend = FernetBackend
sa.PGCryptoBackend = PGCryptoBackend
# revision identifiers, used by Alembic.
revision = '49a9e85a0ffc'
down_revision = 'ed41acf21270'
branch_labels = None
depends_on = None
def upgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
schema_upgrades()
data_upgrades()
def downgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
data_downgrades()
schema_downgrades()
def schema_upgrades() -> None:
"""schema upgrade migrations go here."""
pass
def schema_downgrades() -> None:
"""schema downgrade migrations go here."""
pass
def data_upgrades() -> None:
"""
Rewrite every author into the canonical form, merging the rows that collide.
`Author.name` is only now guaranteed tidy — until this revision extractors wrote
whatever the file said, so one person could hold several rows: "Sam Newman" beside
"Newman, Sam;" beside "Sam Newman.epub" (the last from a filename whose extension
was never stripped). Each showed up as its own author in the sidebar and its own
filter, and no amount of fixing the extractors repairs a row already written.
Rows that canonicalize onto one name are merged into the lowest id, which keeps
whichever row the library has been referring to longest. `book.path` is stored, not
derived, so renaming an author moves nothing on disk.
"""
from chitai.services.matching import format_author_name, normalize_author
connection = op.get_bind()
authors = connection.execute(sa.text("SELECT id, name FROM authors")).fetchall()
groups: dict[str, list[int]] = {}
for id, name in sorted(authors):
# A name with nothing left of it after tidying is left exactly as it was:
# merging those together would invent one author out of several unrelated
# broken rows, which is worse than leaving the mess visible.
if canonical := format_author_name(name):
groups.setdefault(canonical, []).append(id)
for canonical, ids in groups.items():
winner, losers = ids[0], ids[1:]
for loser in losers:
# A book credited to both rows would otherwise breach the
# (book_id, author_id) unique constraint the moment the link is repointed.
connection.execute(
sa.text(
"DELETE FROM book_author_links WHERE author_id = :loser AND book_id IN"
" (SELECT book_id FROM book_author_links WHERE author_id = :winner)"
),
{"loser": loser, "winner": winner},
)
connection.execute(
sa.text(
"UPDATE book_author_links SET author_id = :winner"
" WHERE author_id = :loser"
),
{"loser": loser, "winner": winner},
)
connection.execute(
sa.text("DELETE FROM authors WHERE id = :loser"), {"loser": loser}
)
# Only after the losers are gone, or this collides with the unique index.
connection.execute(
sa.text(
"UPDATE authors SET name = :name, normalized_name = :key WHERE id = :id"
),
{"id": winner, "name": canonical, "key": normalize_author(canonical)},
)
def data_downgrades() -> None:
"""
Nothing to undo.
The rows a merge removed are gone, and the spellings it replaced were never
recorded anywhere else — there is nothing to restore them from.
"""
@@ -0,0 +1,122 @@
"""recompute book matching keys
Revision ID: ed41acf21270
Revises: 4358e7d4743a
Create Date: 2026-08-15 15:44:28.341020
"""
import warnings
import sqlalchemy as sa
from alembic import op
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash, FernetBackend
from advanced_alchemy.types.encrypted_string import PGCryptoBackend
from advanced_alchemy.types.password_hash.argon2 import Argon2Hasher
from advanced_alchemy.types.password_hash.passlib import PasslibHasher
from advanced_alchemy.types.password_hash.pwdlib import PwdlibHasher
from sqlalchemy import Text # noqa: F401
__all__ = ["downgrade", "upgrade", "schema_upgrades", "schema_downgrades", "data_upgrades", "data_downgrades"]
sa.GUID = GUID
sa.DateTimeUTC = DateTimeUTC
sa.ORA_JSONB = ORA_JSONB
sa.EncryptedString = EncryptedString
sa.EncryptedText = EncryptedText
sa.StoredObject = StoredObject
sa.PasswordHash = PasswordHash
sa.Argon2Hasher = Argon2Hasher
sa.PasslibHasher = PasslibHasher
sa.PwdlibHasher = PwdlibHasher
sa.FernetBackend = FernetBackend
sa.PGCryptoBackend = PGCryptoBackend
# revision identifiers, used by Alembic.
revision = 'ed41acf21270'
down_revision = '4358e7d4743a'
branch_labels = None
depends_on = None
def upgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
schema_upgrades()
data_upgrades()
def downgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
data_downgrades()
schema_downgrades()
def schema_upgrades() -> None:
"""schema upgrade migrations go here."""
pass
def schema_downgrades() -> None:
"""schema downgrade migrations go here."""
pass
def data_upgrades() -> None:
"""
Recompute every matching key against the current normalization.
The keys are derived, so changing a helper in `services/matching.py` silently
invalidates every row already written — a book stored under the old rules simply
stops matching one stored under the new ones, with nothing to show that anything
is wrong. `normalize_title` learned to strip the compact edition markers a cover
actually carries ("2E", "5e"), which moved `Building Microservices, 2E` onto the
same key as `Building Microservices`.
Any later change to those helpers wants a revision that looks exactly like this
one. It is idempotent and safe to re-run.
"""
from chitai.services.matching import (
normalize_author,
normalize_identifier,
normalize_title,
)
connection = op.get_bind()
books = connection.execute(sa.text("SELECT id, title FROM books")).fetchall()
_apply(
connection,
"UPDATE books SET normalized_title = :key WHERE id = :id",
[{"id": id, "key": normalize_title(title)} for id, title in books],
)
authors = connection.execute(sa.text("SELECT id, name FROM authors")).fetchall()
_apply(
connection,
"UPDATE authors SET normalized_name = :key WHERE id = :id",
[{"id": id, "key": normalize_author(name)} for id, name in authors],
)
identifiers = connection.execute(
sa.text("SELECT id, name, value FROM identifiers")
).fetchall()
_apply(
connection,
"UPDATE identifiers SET normalized_value = :key WHERE id = :id",
[
{"id": id, "key": normalize_identifier(name, value)}
for id, name, value in identifiers
],
)
def _apply(connection, statement: str, parameters: list[dict]) -> None:
"""Run one update per row, in batches, skipping the work when there are none."""
batch_size = 1000
for start in range(0, len(parameters), batch_size):
connection.execute(sa.text(statement), parameters[start : start + batch_size])
def data_downgrades() -> None:
"""Add any optional data downgrade migrations here!"""
@@ -0,0 +1,114 @@
"""backfill file content types
Revision ID: d2d69065ede3
Revises: 49a9e85a0ffc
Create Date: 2026-08-17 11:37:58.959135
"""
import warnings
from typing import TYPE_CHECKING
import sqlalchemy as sa
from alembic import op
from advanced_alchemy.types import EncryptedString, EncryptedText, GUID, ORA_JSONB, DateTimeUTC, StoredObject, PasswordHash, FernetBackend
from advanced_alchemy.types.encrypted_string import PGCryptoBackend
from advanced_alchemy.types.password_hash.argon2 import Argon2Hasher
from advanced_alchemy.types.password_hash.passlib import PasslibHasher
from advanced_alchemy.types.password_hash.pwdlib import PwdlibHasher
from sqlalchemy import Text # noqa: F401
if TYPE_CHECKING:
from collections.abc import Sequence
__all__ = ["downgrade", "upgrade", "schema_upgrades", "schema_downgrades", "data_upgrades", "data_downgrades"]
sa.GUID = GUID
sa.DateTimeUTC = DateTimeUTC
sa.ORA_JSONB = ORA_JSONB
sa.EncryptedString = EncryptedString
sa.EncryptedText = EncryptedText
sa.StoredObject = StoredObject
sa.PasswordHash = PasswordHash
sa.Argon2Hasher = Argon2Hasher
sa.PasslibHasher = PasslibHasher
sa.PwdlibHasher = PwdlibHasher
sa.FernetBackend = FernetBackend
sa.PGCryptoBackend = PGCryptoBackend
# revision identifiers, used by Alembic.
revision = 'd2d69065ede3'
down_revision = '49a9e85a0ffc'
branch_labels = None
depends_on = None
def upgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
schema_upgrades()
data_upgrades()
def downgrade() -> None:
with warnings.catch_warnings():
warnings.filterwarnings("ignore", category=UserWarning)
with op.get_context().autocommit_block():
data_downgrades()
schema_downgrades()
def schema_upgrades() -> None:
"""schema upgrade migrations go here."""
pass
def schema_downgrades() -> None:
"""schema downgrade migrations go here."""
pass
def data_upgrades() -> None:
"""
Name the format of every file whose content type was never worked out.
`create_many_from_existing_files` filled the column from `mimetypes.guess_type`,
which answers None for `.mobi`, `.azw`, `.fb2` and `.lit` — so a consume-directory
import of any of those stored a null, and the OPDS acquisition link a reader app
uses to decide what it can open carried nothing.
Both write paths now go through `guess_content_type`, which this uses too, so the
formats in its table get named retroactively. A row it still cannot name is **left
null** rather than filled with a placeholder: null is the truth, the column is
nullable, and the one consumer that needs a string substitutes one itself.
Idempotent: it only looks at rows that carry nothing.
"""
from chitai.services.utils import guess_content_type
connection = op.get_bind()
files = connection.execute(
sa.text(
"SELECT id, path FROM file_metadata "
"WHERE content_type IS NULL OR content_type = ''"
)
).fetchall()
parameters = [
{"id": id, "content_type": content_type}
for id, path in files
if (content_type := guess_content_type(path)) is not None
]
if not parameters:
return
batch_size = 1000
for start in range(0, len(parameters), batch_size):
connection.execute(
sa.text(
"UPDATE file_metadata SET content_type = :content_type WHERE id = :id"
),
parameters[start : start + batch_size],
)
def data_downgrades() -> None:
"""Add any optional data downgrade migrations here!"""
+4 -1
View File
@@ -8,7 +8,7 @@ authors = [
]
requires-python = ">=3.13"
dependencies = [
"advanced-alchemy==1.8.0",
"advanced-alchemy>=1.8.0",
"aiofiles>=24.1.0",
"asyncpg>=0.30.0",
"ebooklib>=0.19",
@@ -41,3 +41,6 @@ dev = [
[tool.pytest.ini_options]
asyncio_mode = "auto"
testpaths = ["tests"]
filterwarnings = [
"ignore::jwt.warnings.InsecureKeyLengthWarning",
]
+1 -1
View File
@@ -3,7 +3,6 @@
pkgs.mkShell {
buildInputs = with pkgs; [
# Python development environment for Chitai
python313Full
python313Packages.greenlet
python313Packages.ruff
uv
@@ -11,6 +10,7 @@ pkgs.mkShell {
# postgres database
postgresql
claude-code
];
shellHook = ''
+27 -13
View File
@@ -1,5 +1,6 @@
import asyncio
from typing import Any
from contextlib import asynccontextmanager
from typing import Any, AsyncGenerator
from chitai.services.book import BookService
from chitai.services.consume import ConsumeDirectoryWatcher
@@ -15,6 +16,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from chitai import controllers as c
from chitai.cli import CalibreCLIPlugin
from chitai.config import settings
from chitai.database.config import alchemy
from chitai.database.models.user import User
@@ -63,14 +65,15 @@ oauth2_auth = OAuth2PasswordBearerAuth[User](
"/access/signup",
"/opds",
"/schema",
"/syncs",
"/users/auth",
],
)
watcher_task: asyncio.Task
async def startup():
"""Run setup."""
@asynccontextmanager
async def setup_db_connection(app: Litestar) -> AsyncGenerator[None, None]:
# Setup databse
async with settings.alchemy_config.get_session() as db_session:
# Create default library if none exist
@@ -86,21 +89,30 @@ async def startup():
)
await db_session.commit()
# book_service = BookService(session=db_session)
try:
yield
finally:
await db_session.aclose()
@asynccontextmanager
async def setup_directory_watcher(app: Litestar) -> AsyncGenerator[None, None]:
# Create book covers directory if it does not exist
await create_directory(settings.book_cover_path)
# Create consume directory
await create_directory(settings.consume_path)
async with settings.alchemy_config.get_session() as db_session:
book_service = BookService(session=db_session)
library_service = LibraryService(session=db_session)
# file_watcher = ConsumeDirectoryWatcher(settings.consume_path, library_service, book_service)
# watcher_task = asyncio.create_task(file_watcher.init_watcher())
async def shutdown():
""" Run shutdown tasks. """
file_watcher = ConsumeDirectoryWatcher(settings.consume_path, library_service, book_service)
watcher_task = asyncio.create_task(file_watcher.init_watcher())
try:
yield
finally:
watcher_task.cancel()
def create_app() -> Litestar:
@@ -114,13 +126,15 @@ def create_app() -> Litestar:
c.PublisherController,
c.TagController,
c.OpdsController,
c.DeviceController,
c.KosyncController,
create_static_files_router(path="/covers", directories=["./covers"]),
index,
healthcheck,
],
exception_handlers=exception_handlers,
on_startup=[startup],
plugins=[alchemy],
lifespan=[setup_db_connection, setup_directory_watcher],
plugins=[alchemy, CalibreCLIPlugin()],
on_app_init=[oauth2_auth.on_app_init],
openapi_config=OpenAPIConfig(
title="Chitai",
+186
View File
@@ -0,0 +1,186 @@
# src/chitai/cli.py
"""
Extra commands on the `litestar` CLI.
Registered through `CalibreCLIPlugin` in `app.py`, so they run as
`litestar --app-dir src/chitai/ calibre-import …` and get the app's own configuration
without a second way to load it.
The Calibre import lives here as well as behind an endpoint because the case it exists
for is a one-time migration of a library that may be hundreds of gigabytes. That should
not depend on a browser tab staying open.
"""
from __future__ import annotations
import asyncio
from pathlib import Path
import click
from click import Group
from litestar.plugins import CLIPluginProtocol
from chitai.config import settings
from chitai.database.models import Library
from chitai.services.book import BookService, CalibreImportProgress, CalibreImportResult
from chitai.services.calibre import CalibreLibrary, CalibreLibraryError
from chitai.services.library import LibraryService
class CalibreCLIPlugin(CLIPluginProtocol):
"""Adds `calibre-import` to the Litestar CLI."""
def on_cli_init(self, cli: Group) -> None:
cli.add_command(calibre_import)
@click.command(name="calibre-import")
@click.argument(
"source",
type=click.Path(exists=True, file_okay=False, path_type=Path),
)
@click.option(
"--library",
"library_slug",
required=True,
help="Slug of the Chitai library to import into.",
)
@click.option(
"--allow-duplicates",
is_flag=True,
help="Import books whose files the library already holds.",
)
@click.option(
"--dry-run",
is_flag=True,
help="Read the catalogue and report what it holds, without writing anything.",
)
def calibre_import(
source: Path, library_slug: str, allow_duplicates: bool, dry_run: bool
) -> None:
"""
Import a Calibre library from SOURCE, the directory holding its metadata.db.
Files are copied, never moved: the Calibre library is left exactly as it is, and
re-running skips whatever is already stored.
"""
asyncio.run(_import(source, library_slug, allow_duplicates, dry_run))
async def _import(
source: Path, library_slug: str, allow_duplicates: bool, dry_run: bool
) -> None:
try:
library_source = CalibreLibrary(source)
await library_source.open()
except CalibreLibraryError as exc:
raise click.ClickException(str(exc)) from exc
try:
if dry_run:
await _report(library_source)
return
async with settings.alchemy_config.get_session() as session:
library = await _library(session, library_slug)
result = await BookService(session=session).create_many_from_calibre(
library_source,
library,
allow_duplicates=allow_duplicates,
on_progress=_print_progress,
)
finally:
await library_source.close()
_print_summary(result)
if result.failed:
raise SystemExit(1)
async def _library(session: object, slug: str) -> Library:
"""Resolve the target library, or explain what the options were."""
service = LibraryService(session=session) # type: ignore[arg-type]
library = await service.get_one_or_none(Library.slug == slug)
if library is None:
available = ", ".join(sorted(item.slug for item in await service.list()))
raise click.ClickException(
f"No library with slug '{slug}'. Available: {available or 'none'}"
)
# A read-only library is one pointing at a tree Chitai does not own. Copying books
# into it would write into somebody else's directory.
if library.read_only:
raise click.ClickException(
f"Library '{slug}' is read-only, so nothing can be imported into it"
)
return library
async def _report(source: CalibreLibrary) -> None:
"""Describe the catalogue without touching the database."""
books = await source.books()
click.echo(f"{len(books)} book(s) in {source.root}\n")
for book in books:
authors = ", ".join(book.authors) or "unknown author"
formats = ", ".join(file.format for file in book.files) or "no files"
click.echo(f" #{book.calibre_id:<6} {book.title}")
click.echo(f" {'':<7} {authors} · {formats}")
missing = [
book
for book in books
if any(not file.path.is_file() for file in book.files) or not book.files
]
if missing:
click.echo(
f"\n{len(missing)} book(s) have files the catalogue lists "
"but disk does not:"
)
for book in missing:
click.echo(f" #{book.calibre_id} {book.title}")
def _print_progress(progress: CalibreImportProgress) -> None:
marker = {"created": "+", "skipped": "-", "failed": "!"}.get(progress.outcome, " ")
detail = f" ({progress.detail})" if progress.detail else ""
click.echo(
f"[{progress.processed:>5}/{progress.total}] {marker} {progress.title}{detail}"
)
def _print_summary(result: CalibreImportResult) -> None:
click.echo(
f"\n{len(result.created)} created, {len(result.skipped)} skipped, "
f"{len(result.failed)} failed, of {result.total}."
)
if result.duplicate_files:
click.echo(
f"{len(result.duplicate_files)} individual file(s) were already stored and "
"were left out of books that imported otherwise."
)
for failure in result.failed:
click.echo(f" failed #{failure.calibre_id} {failure.title}: {failure.reason}")
# Imported all the same — a metadata match is a guess, and the duplicates screen is
# where these get decided.
for possible in result.possible_duplicates:
names = ", ".join(
f"{candidate.title} (#{candidate.book_id})"
for candidate in possible.candidates
)
click.echo(
f" possible duplicate {possible.title} (#{possible.book_id}) "
f"may already be in the library as: {names}"
)
+24
View File
@@ -1,9 +1,25 @@
from enum import StrEnum
from pydantic import Field, PostgresDsn, computed_field
from pydantic_settings import BaseSettings, SettingsConfigDict
from advanced_alchemy.extensions.litestar import (
SQLAlchemyAsyncConfig,
)
class DuplicateScope(StrEnum):
"""How widely an incoming file is compared against what is already stored."""
LIBRARY = "library"
"""Only files in the library being uploaded to count as duplicates."""
GLOBAL = "global"
"""A file already held by any library counts as a duplicate."""
OFF = "off"
"""No duplicate detection at all."""
class Settings(BaseSettings):
version: str = Field("0.0.1")
project_name: str = Field("chitai")
@@ -33,6 +49,14 @@ class Settings(BaseSettings):
# Path to consume directory
consume_path: str = Field("./consume")
# Duplicate detection
duplicate_scope: DuplicateScope = Field(DuplicateScope.LIBRARY)
# Where the consume watcher parks files it refused as duplicates. Must sit
# outside `consume_path`, or the watcher picks them straight back up and
# tries to resolve the directory name as a library slug.
duplicate_path: str = Field("./duplicates")
@computed_field
@property
def postgres_uri(self) -> PostgresDsn:
@@ -6,3 +6,5 @@ from .author import AuthorController
from .tag import TagController
from .publisher import PublisherController
from .opds import OpdsController
from .kosync_device import DeviceController
from .kosync_progress import KosyncController
+248 -9
View File
@@ -12,7 +12,12 @@ from litestar.params import Dependency, Body
from litestar.enums import RequestEncodingType
from litestar.response import File, Stream
from litestar.exceptions import HTTPException
from litestar.status_codes import HTTP_400_BAD_REQUEST
from litestar.status_codes import (
HTTP_200_OK,
HTTP_204_NO_CONTENT,
HTTP_400_BAD_REQUEST,
HTTP_409_CONFLICT,
)
from litestar.datastructures import UploadFile
from advanced_alchemy.service.pagination import OffsetPagination
from advanced_alchemy.filters import CollectionFilter
@@ -23,6 +28,24 @@ from chitai.services import dependencies as deps
from chitai import schemas as s
from chitai.database import models as m
from chitai.services import BookService, BookProgressService
from chitai.services.book import DuplicateFilesError
def _duplicate_conflict(exc: DuplicateFilesError) -> HTTPException:
"""
Turn refused files into a 409 the caller can act on.
The files ride along in `extra` so the client can name them and offer to send them
again with `allow_duplicates`, rather than being told only that something clashed.
"""
return HTTPException(
status_code=HTTP_409_CONFLICT,
detail="These files are already in the library",
extra=[
s.DuplicateFileRead.model_validate(duplicate).model_dump(mode="json")
for duplicate in exc.duplicates
],
)
class BookController(Controller):
@@ -63,6 +86,7 @@ class BookController(Controller):
books_service: BookService,
library: m.Library,
data: Annotated[s.BookCreate, Body(media_type=RequestEncodingType.MULTI_PART)],
allow_duplicates: bool = False,
) -> s.BookRead:
"""
Create a new book with metadata and files.
@@ -73,6 +97,10 @@ class BookController(Controller):
Path Parameters:
library_id: The ID of the library the book belongs to.
Query Parameters:
allow_duplicates: If True, store the files even if the library already
holds them.
Request Body:
data: Book creation data including metadata and files.
@@ -83,9 +111,17 @@ class BookController(Controller):
Returns:
The created book as a BookRead schema.
Raises:
HTTPException: 409 if any of the files is already in the library.
"""
result = await books_service.create_book(data, library)
try:
result = await books_service.create_book(
data, library, screen_duplicates=not allow_duplicates
)
except DuplicateFilesError as exc:
raise _duplicate_conflict(exc)
book = await books_service.get(result.id)
return books_service.to_schema(book, schema_type=s.BookRead)
@@ -97,13 +133,20 @@ class BookController(Controller):
data: Annotated[
s.BooksCreateFromFiles, Body(media_type=RequestEncodingType.MULTI_PART)
],
) -> OffsetPagination[s.BookRead]:
allow_duplicates: bool = False,
) -> s.BooksUploadResult:
"""
Create multiple books from uploaded files.
Groups files by directory and creates separate books for each group.
Metadata is automatically extracted from the files.
Files the library already holds are skipped rather than refused, and reported
back so the caller can say which ones did not make it in and why.
Query Parameters:
allow_duplicates: If True, store every file, even one already held.
Request Body:
data: Container with list of uploaded files.
@@ -112,19 +155,200 @@ class BookController(Controller):
library: The library the books belong to.
Returns:
Paginated list of created books.
The books created, and the files skipped as duplicates.
"""
try:
results = await books_service.create_many_from_files(data, library)
result = await books_service.create_many_from_files(
data, library, allow_duplicates=allow_duplicates
)
except ValueError:
raise HTTPException(
status_code=HTTP_400_BAD_REQUEST, detail="Must upload at least one file"
)
books = await books_service.list(
CollectionFilter("id", [result.id for result in results])
books = (
await books_service.list(
CollectionFilter("id", [book.id for book in result.books])
)
return books_service.to_schema(books, schema_type=s.BookRead)
if result.books
else []
)
return s.BooksUploadResult(
created=[
books_service.to_schema(book, schema_type=s.BookRead) for book in books
],
skipped=[
s.DuplicateFileRead.model_validate(duplicate)
for duplicate in result.duplicates
],
possible_duplicates=[
s.PossibleDuplicateRead.model_validate(possible)
for possible in result.possible_duplicates
],
)
@get(path="duplicate-books")
async def list_duplicate_books(
self, books_service: BookService, library: m.Library
) -> list[s.DuplicateBookGroupRead]:
"""
Report books already in the library that look like copies of one another.
The import-time check only ever sees what is arriving, so this is what covers
a collection someone already has. Matching is on metadata and therefore a
guess: a group is a question for the reader, not a verdict.
Query Parameters:
library_id: The library to review.
Injected Dependencies:
books_service: The book service for database operations.
library: The library to review.
Returns:
One entry per group of two or more books. Empty when there is nothing to
review, or when duplicate detection is switched off.
"""
groups = await books_service.find_duplicate_book_groups(library)
return [
s.DuplicateBookGroupRead(
books=[s.DuplicateBookRead.model_validate(book) for book in group]
)
for group in groups
]
@post(path="merge")
async def merge_books(
self, books_service: BookService, library: m.Library, data: s.BookMerge
) -> s.BookRead:
"""
Fold several books into one and delete the records folded in.
The survivor keeps its id, so links and bookmarks still resolve. Files, reading
progress, shelves, tags and unheld identifiers move onto it; metadata is only
changed by what `metadata` names, because choosing between two titles is the
reader's judgement rather than this endpoint's.
Nothing is removed from disk — a wrong merge should cost metadata that can be
retyped, not a book.
Query Parameters:
library_id: The library the books belong to.
Request Body:
data: The survivor, the books to fold in, and the resolved metadata.
Injected Dependencies:
books_service: The book service for database operations.
library: The library the books belong to.
Returns:
The surviving book.
Raises:
HTTPException: 400 if fewer than two distinct books were named, one is
unknown, or they do not all belong to one library.
"""
try:
book = await books_service.merge_books(
data.survivor_id,
data.merged_ids,
library,
metadata=data.metadata.model_dump(exclude_unset=True)
if data.metadata
else None,
)
except ValueError as exc:
raise HTTPException(status_code=HTTP_400_BAD_REQUEST, detail=str(exc))
return books_service.to_schema(book, schema_type=s.BookRead)
@post(path="duplicate-books/dismissals", status_code=HTTP_204_NO_CONTENT)
async def dismiss_duplicate_books(
self, books_service: BookService, data: s.DuplicateDismissal
) -> None:
"""
Record that two books are not the same book.
Without this the review screen proposes the same wrong pair forever, which is
how a reader learns to stop looking at it.
Request Body:
data: The two book IDs. Order does not matter.
Injected Dependencies:
books_service: The book service for database operations.
Raises:
HTTPException: 400 if the two IDs are the same or either book is unknown.
"""
try:
await books_service.dismiss_duplicates(data.book_a_id, data.book_b_id)
except ValueError as exc:
raise HTTPException(status_code=HTTP_400_BAD_REQUEST, detail=str(exc))
@delete(path="duplicate-books/dismissals")
async def restore_duplicate_books(
self, books_service: BookService, book_a_id: int, book_b_id: int
) -> None:
"""
Undo a dismissal, so the pair is proposed again.
Query Parameters:
book_a_id: One of the two books.
book_b_id: The other. Order does not matter.
Injected Dependencies:
books_service: The book service for database operations.
"""
await books_service.restore_duplicates(book_a_id, book_b_id)
# A question, not a change: 200 rather than the 201 a POST would default to.
@post(path="duplicate-files", status_code=HTTP_200_OK)
async def check_duplicate_files(
self,
books_service: BookService,
library: m.Library,
data: list[s.FileFingerprint],
) -> list[s.DuplicateFileRead]:
"""
Report which of the given files the library already holds.
Lets a client ask before it uploads anything, which is the difference between
re-sending a folder of books and re-sending twelve kilobytes of hashes.
Query Parameters:
library_id: The library to check against.
Request Body:
data: Hash and size for each file, optionally with the name to echo back.
Injected Dependencies:
books_service: The book service for database operations.
library: The library to check against.
Returns:
One entry per submitted file that is already stored. Files that are not
are absent.
"""
matches = await books_service.find_duplicate_files(
((item.hash, item.size) for item in data), library
)
return [
s.DuplicateFileRead(
filename=item.filename or match.filename,
hash=match.hash,
size=match.size,
library_id=match.library_id,
book_id=match.book_id,
book_title=match.book_title,
)
for item in data
if (match := matches.get((item.hash, item.size))) is not None
]
@get(path="/{book_id:int}")
async def get_book_by_id(
@@ -303,13 +527,20 @@ class BookController(Controller):
],
library: m.Library,
books_service: BookService,
allow_duplicates: bool = False,
) -> s.BookRead:
"""
Add files to an existing book.
A file the book already carries is ignored, so re-sending one is harmless.
Path Parameters:
book_id: The ID of the book to modify
Query Parameters:
allow_duplicates: If True, store the files even if the library already
holds them.
Request Body:
files: The files to add to the book
@@ -320,9 +551,17 @@ class BookController(Controller):
Returns:
The modified book
Raises:
HTTPException: 409 if a file is already stored under a different book.
"""
await books_service.add_files(book_id, data, library)
try:
await books_service.add_files(
book_id, data, library, allow_duplicates=allow_duplicates
)
except DuplicateFilesError as exc:
raise _duplicate_conflict(exc)
book = await books_service.get(book_id)
return books_service.to_schema(book, schema_type=s.BookRead)
@@ -0,0 +1,54 @@
from advanced_alchemy.service.pagination import OffsetPagination
from chitai.database.models.kosync_device import KosyncDevice
from chitai.database.models.user import User
from chitai.schemas.kosync import KosyncDeviceCreate, KosyncDeviceRead
from chitai.services.kosync_device import KosyncDeviceService
from litestar import Controller, post, get, delete
from litestar.di import Provide
from chitai.services import dependencies as deps
class DeviceController(Controller):
""" Controller for managing KOReader devices."""
dependencies = {
"device_service": Provide(deps.provide_kosync_device_service)
}
path = "/devices"
@get()
async def get_devices(self, device_service: KosyncDeviceService, current_user: User) -> OffsetPagination[KosyncDeviceRead]:
""" Return a list of all the user's devices."""
devices = await device_service.list(
KosyncDevice.user_id == current_user.id
)
return device_service.to_schema(devices, schema_type=KosyncDeviceRead)
@post()
async def create_device(self, data: KosyncDeviceCreate, device_service: KosyncDeviceService, current_user: User) -> KosyncDeviceRead:
device = await device_service.create({
'name': data.name,
'user_id': current_user.id
})
return device_service.to_schema(device, schema_type=KosyncDeviceRead)
@delete("/{device_id:int}")
async def delete_device(self, device_id: int, device_service: KosyncDeviceService, current_user: User) -> None:
# Ensure the device exists and is owned by the user
device = await device_service.get_one(
KosyncDevice.id == device_id,
KosyncDevice.user_id == current_user.id
)
await device_service.delete(device.id)
@get("/{device_id:int}/regenerate")
async def regenerate_device_api_key(self, device_id: int, device_service: KosyncDeviceService, current_user: User) -> KosyncDeviceRead:
# Ensure the device exists and is owned by the user
device = await device_service.get_one(
KosyncDevice.id == device_id,
KosyncDevice.user_id == current_user.id
)
updated_device = await device_service.regenerate_api_key(device.id)
return device_service.to_schema(updated_device, schema_type=KosyncDeviceRead)
@@ -0,0 +1,90 @@
from __future__ import annotations
from typing import Annotated
from litestar import Controller, get, put
from litestar.exceptions import HTTPException
from litestar.status_codes import HTTP_403_FORBIDDEN
from litestar.response import Response
from litestar.params import Parameter
from litestar.di import Provide
from chitai.database import models as m
from chitai.schemas.kosync import KosyncProgressUpdate, KosyncProgressRead
from chitai.services.book import BookService
from chitai.services.kosync_progress import KosyncProgressService
from chitai.services.filters.book import FileHashFilter
from chitai.services import dependencies as deps
from chitai.middleware.kosync_auth import kosync_api_key_auth
class KosyncController(Controller):
"""Controller for syncing progress with KOReader devices."""
middleware = [kosync_api_key_auth]
dependencies = {
"kosync_progress_service": Provide(deps.provide_kosync_progress_service),
"book_service": Provide(deps.provide_book_service),
"user": Provide(deps.provide_user_via_kosync_auth),
}
@put("/syncs/progress")
async def upload_progress(
self,
data: KosyncProgressUpdate,
book_service: BookService,
kosync_progress_service: KosyncProgressService,
user: m.User,
) -> None:
"""Upload book progress from a KOReader device."""
book = await book_service.get_one(FileHashFilter([data.document]))
await kosync_progress_service.upsert_progress(
user_id=user.id,
book_id=book.id,
document=data.document,
progress=data.progress,
percentage=data.percentage,
device=data.device,
device_id=data.device_id,
)
@get("/syncs/progress/{document_id:str}")
async def get_progress(
self,
document_id: str,
kosync_progress_service: KosyncProgressService,
user: m.User,
) -> KosyncProgressRead:
"""Return the Kosync progress record associated with the given document."""
progress = await kosync_progress_service.get_by_document_hash(user.id, document_id)
if not progress:
raise HTTPException(status_code=404, detail="No progress found for document")
return KosyncProgressRead(
document=progress.document,
progress=progress.progress,
percentage=progress.percentage,
device=progress.device,
device_id=progress.device_id,
)
@get("/users/auth")
async def authorize(
self, _api_key: Annotated[str, Parameter(header="X-AUTH-USER")]
) -> Response[dict[str, str]]:
"""Verify authentication (handled by middleware)."""
return Response(status_code=200, content={"authorized": "OK"})
@get("/users/register")
async def register(self) -> None:
"""User registration endpoint - disabled."""
raise HTTPException(
detail="User accounts must be created via the main application",
status_code=HTTP_403_FORBIDDEN,
)
+165 -2
View File
@@ -1,12 +1,20 @@
# src/chitai/controllers/library.py
# Standard library
import asyncio
import shutil
import tempfile
from pathlib import Path
from typing import Annotated
# Third-party libraries
import aiofiles
from aiofiles import os as aios
from litestar import Controller, post, get, patch, delete
from litestar.params import Dependency
from litestar.enums import RequestEncodingType
from litestar.params import Body, Dependency
from litestar.exceptions import HTTPException
from litestar.status_codes import HTTP_200_OK, HTTP_202_ACCEPTED
from advanced_alchemy.extensions.litestar.providers import create_service_dependencies
from advanced_alchemy.service.pagination import OffsetPagination
from advanced_alchemy.service import FilterTypeT
@@ -14,10 +22,21 @@ from advanced_alchemy.service import FilterTypeT
# Local imports
from chitai.database import models as m
from chitai.services import LibraryService
from chitai.schemas.library import LibraryCreate, LibraryRead
from chitai.schemas.library import (
CalibreArchiveUpload,
CalibreImportRead,
LibraryCreate,
LibraryRead,
)
from chitai.services.calibre import CalibreLibraryError, extract_calibre_archive
from chitai.services.calibre_import import registry
from chitai.services.utils import DirectoryDoesNotExist
# How much of an uploaded archive is held in memory at a time on its way to disk.
UPLOAD_CHUNK_SIZE = 262144 # 256 KiB
class LibraryController(Controller):
"""Controller for managing library operations."""
@@ -74,3 +93,147 @@ class LibraryController(Controller):
return library_service.to_schema(
results, total, filters, schema_type=LibraryRead
)
@post(
path="{library_id:int}/imports/calibre/upload",
status_code=HTTP_202_ACCEPTED,
request_max_body_size=None,
)
async def upload_calibre_import(
self,
library_service: LibraryService,
library_id: int,
data: Annotated[
CalibreArchiveUpload, Body(media_type=RequestEncodingType.MULTI_PART)
],
) -> CalibreImportRead:
"""
Import a zipped Calibre library that was uploaded rather than named on disk.
For the case where the library is not on the server: zip the Calibre folder and
send it. Unpacked into a temp directory the job owns and deletes when it ends —
by which time the books worth keeping have been copied into the library proper.
The server-folder route stays the one for a very large library. This one has to
carry the whole archive over HTTP first.
Path Parameters:
library_id: The library to import into.
Request Body:
data: The `.zip` holding the Calibre library.
Returns:
The job, already running.
Raises:
HTTPException: 400 if the archive is not a zip, holds no `metadata.db`,
names an entry outside itself, or would not fit on disk; 409 if an
import into this library is already running.
"""
library = await self._importable(library_service, library_id)
if (running := registry.running_for(library_id)) is not None:
raise HTTPException(
status_code=409,
detail="An import into this library is already running",
extra={"job_id": running.id},
)
workspace = Path(await asyncio.to_thread(tempfile.mkdtemp))
try:
archive = workspace / "upload.zip"
await data.archive.seek(0)
async with aiofiles.open(archive, "wb") as destination:
while chunk := await data.archive.read(UPLOAD_CHUNK_SIZE):
await destination.write(chunk)
unpacked = workspace / "library"
unpacked.mkdir()
catalogue = await extract_calibre_archive(archive, unpacked)
# The archive itself is dead weight once unpacked, and the library it
# unpacked to can be large.
await aios.remove(archive)
except CalibreLibraryError as exc:
await asyncio.to_thread(shutil.rmtree, workspace, True)
raise HTTPException(status_code=400, detail=str(exc))
except Exception:
await asyncio.to_thread(shutil.rmtree, workspace, True)
raise
job = registry.start(
library,
catalogue,
workspace=workspace,
label=data.archive.filename or "uploaded archive",
allow_duplicates=data.allow_duplicates,
)
return CalibreImportRead.model_validate(job)
@get(path="imports/{job_id:str}")
async def get_import(self, job_id: str) -> CalibreImportRead:
"""
Report on an import.
Polled by the client while a run is going. Jobs are held in memory, so this is
answered by the process that started it — see `services/calibre_import.py`.
Path Parameters:
job_id: The job to report on.
Raises:
HTTPException: 404 if this process holds no such job.
"""
if (job := registry.get(job_id)) is None:
raise HTTPException(status_code=404, detail="No such import")
return CalibreImportRead.model_validate(job)
@delete(path="imports/{job_id:str}", status_code=HTTP_200_OK)
async def cancel_import(self, job_id: str) -> CalibreImportRead:
"""
Ask an import to stop after the book it is on.
Deliberately not an abort: a book abandoned mid-copy would leave files on disk
with no row describing them. Whatever it has imported stays imported.
Path Parameters:
job_id: The job to stop.
Raises:
HTTPException: 404 if this process holds no such job.
"""
if (job := registry.cancel(job_id)) is None:
raise HTTPException(status_code=404, detail="No such import")
return CalibreImportRead.model_validate(job)
@staticmethod
async def _importable(
library_service: LibraryService, library_id: int
) -> m.Library:
"""
The library, if it can be imported into at all.
Raises:
HTTPException: 404 if there is no such library, 400 if it is read-only —
a read-only library points at a tree Chitai does not own, so copying
books into it would write into somebody else's directory.
"""
library = await library_service.get_one_or_none(m.Library.id == library_id)
if library is None:
raise HTTPException(status_code=404, detail="No such library")
if library.read_only:
raise HTTPException(
status_code=400,
detail="This library is read-only, so nothing can be imported into it",
)
return library
+4 -3
View File
@@ -5,8 +5,9 @@ from advanced_alchemy.extensions.litestar import (
AsyncSessionConfig,
SQLAlchemyPlugin,
)
from advanced_alchemy.base import BigIntAuditBase
from sqlalchemy.ext.asyncio import create_async_engine
from chitai.database import models
from chitai.database import models # noqa: F401 # Import to register models
DATABASE_URL = str(settings.postgres_uri)
@@ -16,8 +17,8 @@ config = SQLAlchemyAsyncConfig(
engine_instance=create_async_engine(DATABASE_URL, echo=settings.postgres_echo),
session_config=session_config,
before_send_handler="autocommit",
create_all=True,
create_all=False,
metadata=BigIntAuditBase.registry.metadata,
)
alchemy = SQLAlchemyPlugin(config=config)
@@ -3,6 +3,9 @@ from .book import Book, Identifier, FileMetadata
from .book_list import BookList, BookListLink
from .book_progress import BookProgress
from .book_series import BookSeries
from .duplicate_dismissal import DuplicateDismissal
from .kosync_device import KosyncDevice
from .kosync_progress import KosyncProgress
from .library import Library
from .publisher import Publisher
from .tag import Tag, BookTagLink
+44 -7
View File
@@ -1,10 +1,11 @@
from typing import TYPE_CHECKING, Optional
from collections.abc import Hashable
from sqlalchemy import ColumnElement, ForeignKey
from sqlalchemy import ColumnElement, ForeignKey, UniqueConstraint
from sqlalchemy.orm import Mapped
from sqlalchemy.orm import mapped_column
from sqlalchemy.orm import relationship
from sqlalchemy.orm import validates
from advanced_alchemy.base import BigIntAuditBase, BigIntBase
from advanced_alchemy.mixins import UniqueMixin
@@ -16,18 +17,55 @@ if TYPE_CHECKING:
class Author(BigIntAuditBase, UniqueMixin):
__tablename__ = "authors"
# Always the canonical form — see `_canonicalize`. Extractors hand over whatever
# the file happened to say: "Newman, Sam;" from a `DC:creator` list, or
# "Sam Newman.epub" from a filename. Storing those verbatim is how one person ends
# up as several rows in the sidebar.
name: Mapped[str] = mapped_column(unique=True, index=True)
# Kept current by `_canonicalize` too — never assign it directly. Not unique: two
# spellings that survive canonicalization, "Steve Mcconnell" and "Steve McConnell",
# are still one person to a reader, which is what this column exists to express.
normalized_name: Mapped[str] = mapped_column(default="", index=True)
description: Mapped[Optional[str]]
@validates("name")
def _canonicalize(self, _key: str, name: str) -> str:
"""
Store the tidied name, and derive the matching key from it.
A validator so no write can get around it, and `unique_hash` / `unique_filter`
below tidy the same way so `as_unique_async` looks the row up under the name it
would actually be stored as. All three have to agree: if the lookup used the
raw name and the insert used the tidy one, every variant spelling would miss
the existing row and then collide with it on the unique index.
"""
# Imported here rather than at module scope: `chitai.services.matching` cannot
# be reached without initialising the `chitai.services` package, which imports
# the services, which import this module.
from chitai.services.matching import format_author_name, normalize_author
name = format_author_name(name)
self.normalized_name = normalize_author(name)
return name
@classmethod
def _tidy(cls, name: str) -> str:
"""The name as `_canonicalize` would store it."""
from chitai.services.matching import format_author_name
return format_author_name(name)
@classmethod
def unique_hash(cls, name: str) -> Hashable:
"""Generate a unique hash for deduplication."""
return name
return cls._tidy(name)
@classmethod
def unique_filter(cls, name: str) -> ColumnElement[bool]:
"""SQL filter for finding existing records."""
return cls.name == name
return cls.name == cls._tidy(name)
def __repr__(self) -> str:
return f"Author({self.name!r})"
@@ -35,11 +73,10 @@ class Author(BigIntAuditBase, UniqueMixin):
class BookAuthorLink(BigIntBase):
__tablename__ = "book_author_links"
__table_args__ = (UniqueConstraint("book_id", "author_id"),)
book_id: Mapped[int] = mapped_column(
ForeignKey("books.id", ondelete="cascade"), primary_key=True
)
author_id: Mapped[int] = mapped_column(ForeignKey("authors.id"), primary_key=True)
book_id: Mapped[int] = mapped_column(ForeignKey("books.id", ondelete="cascade"))
author_id: Mapped[int] = mapped_column(ForeignKey("authors.id"))
position: Mapped[int]
+57 -10
View File
@@ -1,15 +1,11 @@
from datetime import date
from typing import TYPE_CHECKING, Any, Optional
from sqlalchemy import Index
from sqlalchemy import ForeignKey
from sqlalchemy.orm import Mapped, mapped_column, relationship
from sqlalchemy.orm import mapped_column
from sqlalchemy.orm import relationship
from sqlalchemy import Index, ForeignKey, UniqueConstraint
from sqlalchemy.orm import Mapped, mapped_column, relationship, validates
from sqlalchemy.ext.orderinglist import ordering_list
from sqlalchemy.ext.associationproxy import association_proxy
from sqlalchemy.ext.associationproxy import AssociationProxy
from sqlalchemy.orm.collections import attribute_keyed_dict
from advanced_alchemy.base import BigIntAuditBase, BigIntBase
@@ -45,6 +41,13 @@ class Book(BigIntAuditBase):
library: Mapped["Library"] = relationship(back_populates="books")
title: Mapped[str]
# Kept current by `_normalize_title` below — never assign it directly.
#
# Deliberately not unique: two spellings collapsing onto one value is the whole
# point of the column, and a second edition is allowed to exist.
normalized_title: Mapped[str] = mapped_column(default="", index=True)
subtitle: Mapped[Optional[str]]
description: Mapped[Optional[str]]
published_date: Mapped[Optional[date]]
@@ -112,6 +115,24 @@ class Book(BigIntAuditBase):
def progress(self) -> Optional["BookProgress"]:
return self.progress_records[0] if self.progress_records else None
@validates("title")
def _normalize_title(self, _key: str, title: str) -> str:
"""
Derive `normalized_title` from whatever writes the title.
A validator rather than a service call because `BookService` sets titles from
at least three places — `to_model_on_create`, `to_model_on_update` and the
`setattr` loop in `_populate_with_unique_relationships` — and a fourth would
otherwise leave the key silently stale.
"""
# Imported here rather than at module scope: `chitai.services.matching` cannot
# be reached without initialising the `chitai.services` package, which imports
# the services, which import this module.
from chitai.services.matching import normalize_title
self.normalized_title = normalize_title(title)
return title
def __repr__(self) -> str:
return f"Book({self.title=!r})"
@@ -127,13 +148,30 @@ class Book(BigIntAuditBase):
class Identifier(BigIntBase):
__tablename__ = "identifiers"
__table_args__ = (UniqueConstraint("name", "book_id"),)
name: Mapped[str] = mapped_column(primary_key=True)
book_id: Mapped[int] = mapped_column(
ForeignKey("books.id", ondelete="cascade"), primary_key=True
)
name: Mapped[str]
book_id: Mapped[int] = mapped_column(ForeignKey("books.id", ondelete="cascade"))
value: Mapped[str]
# Kept current by `_normalize` below — never assign it directly. Null for an
# identifier that cannot carry a match: a per-build UUID, or an ISBN that fails
# its own checksum.
normalized_value: Mapped[Optional[str]] = mapped_column(index=True)
@validates("name", "value")
def _normalize(self, key: str, value: str) -> str:
"""Recompute `normalized_value` whenever either half of the pair changes."""
from chitai.services.matching import normalize_identifier
name = value if key == "name" else self.name
raw = value if key == "value" else self.value
self.normalized_value = (
normalize_identifier(name, raw) if name and raw else None
)
return value
def __repr__(self):
return f"Identifier({self.name!r} : {self.value!r})"
@@ -141,6 +179,15 @@ class Identifier(BigIntBase):
class FileMetadata(BigIntBase):
__tablename__ = "file_metadata"
__table_args__ = (
# Deliberately not unique. The hash is KOReader's partial MD5, which samples
# 12 KiB of the file, so two genuinely different files can collide — and an
# existing database may already hold duplicates, which a unique index would
# refuse to build over. Duplicate detection pairs it with `size` and treats a
# match as advisory, so this only has to make the lookup cheap.
Index("ix_file_metadata_hash", "hash"),
)
book_id: Mapped[int] = mapped_column(ForeignKey("books.id", ondelete="cascade"))
book: Mapped[Book] = relationship(back_populates="files")
hash: Mapped[str]
@@ -1,5 +1,5 @@
from typing import Optional
from sqlalchemy import ForeignKey
from sqlalchemy import ForeignKey, UniqueConstraint
from sqlalchemy.orm import Mapped, mapped_column, relationship
from sqlalchemy.ext.associationproxy import association_proxy, AssociationProxy
from sqlalchemy.ext.orderinglist import ordering_list
@@ -38,10 +38,10 @@ class BookList(BigIntAuditBase):
class BookListLink(BigIntBase):
__tablename__ = "book_list_links"
book_id: Mapped[int] = mapped_column(
ForeignKey("books.id", ondelete="cascade"), primary_key=True
)
list_id: Mapped[int] = mapped_column(ForeignKey("book_lists.id"), primary_key=True)
__table_args__ = (UniqueConstraint("book_id", "list_id"),)
book_id: Mapped[int] = mapped_column(ForeignKey("books.id", ondelete="cascade"))
list_id: Mapped[int] = mapped_column(ForeignKey("book_lists.id"))
position: Mapped[int]
book: Mapped[Book] = relationship(back_populates="list_links")
@@ -0,0 +1,37 @@
from sqlalchemy import ForeignKey, UniqueConstraint
from sqlalchemy.orm import Mapped, mapped_column
from advanced_alchemy.base import BigIntBase
class DuplicateDismissal(BigIntBase):
"""
Two books a reader has said are not the same book.
Title and author matching is probabilistic, so it will keep proposing a second
edition, a translation and a sequel that shares its predecessor's name. A review
screen with no way to disagree with it nags forever, which is how people learn to
ignore a screen.
The pair is stored ordered — `book_a_id` is always the lower id — so "A and B" and
"B and A" are one row and the unique constraint can do its job. Use `pair()` rather
than assigning the columns directly.
"""
__tablename__ = "duplicate_dismissals"
__table_args__ = (UniqueConstraint("book_a_id", "book_b_id"),)
book_a_id: Mapped[int] = mapped_column(
ForeignKey("books.id", ondelete="cascade"), index=True
)
book_b_id: Mapped[int] = mapped_column(
ForeignKey("books.id", ondelete="cascade"), index=True
)
@staticmethod
def pair(first: int, second: int) -> tuple[int, int]:
"""The two book ids in the order this table stores them."""
return (first, second) if first <= second else (second, first)
def __repr__(self) -> str:
return f"DuplicateDismissal({self.book_a_id!r}, {self.book_b_id!r})"
@@ -0,0 +1,15 @@
from sqlalchemy import ColumnElement, ForeignKey
from sqlalchemy.orm import Mapped
from sqlalchemy.orm import mapped_column
from advanced_alchemy.base import BigIntAuditBase
class KosyncDevice(BigIntAuditBase):
__tablename__ = "devices"
user_id: Mapped[int] = mapped_column(ForeignKey("users.id"), nullable=False)
api_key: Mapped[str] = mapped_column(unique=True)
name: Mapped[str]
def __repr__(self) -> str:
return f"KosyncDevice({self.name!r})"
@@ -0,0 +1,26 @@
from __future__ import annotations
from typing import Optional
from sqlalchemy import ForeignKey
from sqlalchemy.orm import Mapped, mapped_column
from advanced_alchemy.base import BigIntAuditBase
class KosyncProgress(BigIntAuditBase):
"""Progress tracking for KOReader devices, keyed by document hash."""
__tablename__ = "kosync_progress"
user_id: Mapped[int] = mapped_column(
ForeignKey("users.id", ondelete="cascade"), nullable=False
)
book_id: Mapped[int] = mapped_column(
ForeignKey("books.id", ondelete="cascade"), nullable=False
)
document: Mapped[str] = mapped_column(nullable=False)
progress: Mapped[Optional[str]]
percentage: Mapped[Optional[float]]
device: Mapped[Optional[str]]
device_id: Mapped[Optional[str]]
@@ -17,6 +17,7 @@ class Library(BigIntAuditBase, SlugKey):
# Which structure to save the files in the filesystem (i.e {author_name}/{title}.{ext})
path_template: Mapped[str]
description: Mapped[Optional[str]]
icon: Mapped[str] = mapped_column(default="library")
read_only: Mapped[bool] = mapped_column(nullable=False, default=False)
books: Mapped[list["Book"]] = relationship(back_populates="library")
+4 -7
View File
@@ -1,7 +1,7 @@
from collections.abc import Hashable
from typing import TYPE_CHECKING
from sqlalchemy import ColumnElement, ForeignKey
from sqlalchemy import ColumnElement, ForeignKey, UniqueConstraint
from sqlalchemy.orm import Mapped
from sqlalchemy.orm import mapped_column
from sqlalchemy.orm import relationship
@@ -35,15 +35,12 @@ class Tag(BigIntBase, UniqueMixin):
class BookTagLink(BigIntBase):
__tablename__ = "book_tag_link"
__table_args__ = (UniqueConstraint("book_id", "tag_id"),)
book_id: Mapped[int] = mapped_column(
ForeignKey("books.id", ondelete="cascade"), primary_key=True
)
tag_id: Mapped[int] = mapped_column(ForeignKey("tags.id"), primary_key=True)
book_id: Mapped[int] = mapped_column(ForeignKey("books.id", ondelete="cascade"))
tag_id: Mapped[int] = mapped_column(ForeignKey("tags.id"))
position: Mapped[int]
book: Mapped["Book"] = relationship(back_populates="tag_links")
tag: Mapped[Tag] = relationship()
@@ -0,0 +1,37 @@
from chitai.services.user import UserService
from chitai.services.kosync_device import KosyncDeviceService
from litestar.middleware import (
AbstractAuthenticationMiddleware,
AuthenticationResult,
DefineMiddleware
)
from litestar.connection import ASGIConnection
from litestar.exceptions import NotAuthorizedException, PermissionDeniedException
from chitai.config import settings
class KosyncAuthenticationMiddleware(AbstractAuthenticationMiddleware):
async def authenticate_request(self, connection: ASGIConnection) -> AuthenticationResult:
"""Given a request, parse the header for Base64 encoded basic auth credentials. """
# retrieve the auth header
api_key = connection.headers.get("X-AUTH-USER", None)
if not api_key:
raise NotAuthorizedException()
try:
db_session = settings.alchemy_config.provide_session(connection.app.state, connection.scope)
user_service = UserService(db_session)
device_service = KosyncDeviceService(db_session)
device = await device_service.get_by_api_key(api_key)
user = await user_service.get(device.user_id)
return AuthenticationResult(user=user, auth=None)
except PermissionDeniedException as exc:
print(exc)
raise NotAuthorizedException()
kosync_api_key_auth = DefineMiddleware(KosyncAuthenticationMiddleware)
+8
View File
@@ -4,7 +4,15 @@ from .book import (
BookProgressCreate,
BookProgressRead,
BooksCreateFromFiles,
BooksUploadResult,
BookMerge,
BookMetadataUpdate,
DuplicateBookGroupRead,
DuplicateBookRead,
DuplicateDismissal,
DuplicateFileRead,
FileFingerprint,
PossibleDuplicateRead,
FileMetadataRead,
BookSeriesRead,
)
+106 -9
View File
@@ -30,7 +30,11 @@ class FileMetadataRead(BaseModel):
path: str
hash: str
size: int
content_type: str
# Nullable, though every ingest path now writes one through
# `guess_content_type`. Rows predating it can hold null, and a required field here
# turns one of those into a 500 on a book the reader can otherwise open.
content_type: str | None = None
@computed_field
@property
@@ -38,6 +42,16 @@ class FileMetadataRead(BaseModel):
return Path(self.path).name
class BookProgressRead(BaseModel):
percentage: float
epub_cfi: str | None = None
epub_xpointer: str | None = None
pdf_page: int | None = None
completed: bool | None = False
device_type: str | None = None
device_id: str | None = None
class BookRead(BaseModel):
id: int
library_id: int
@@ -119,6 +133,83 @@ class BooksCreateFromFiles(BaseModel):
model_config = ConfigDict(arbitrary_types_allowed=True)
class FileFingerprint(BaseModel):
"""What a client can say about a file it has not uploaded yet."""
hash: str
size: int
filename: str = ""
class DuplicateFileRead(BaseModel):
"""A file that was not stored because the library already holds its bytes."""
model_config = ConfigDict(from_attributes=True)
filename: str
hash: str
size: int
library_id: int
# Null when the match was another file in the same upload, which has no row yet.
book_id: int | None = None
book_title: str | None = None
class DuplicateBookRead(BaseModel):
"""
A stored book that may be the same book as another one.
Unlike `DuplicateFileRead` this is a guess: the evidence is metadata two editions
of one work legitimately share. Nothing was refused on the strength of it.
"""
model_config = ConfigDict(from_attributes=True)
book_id: int
title: str
authors: list[str]
library_id: int
cover_image: Path | None = None
# Why it matched: "identifier" and/or "title-author".
matched_on: list[str]
class PossibleDuplicateRead(BaseModel):
"""A book that was imported, together with what it might be a second copy of."""
model_config = ConfigDict(from_attributes=True)
book_id: int
title: str
candidates: list[DuplicateBookRead]
class DuplicateBookGroupRead(BaseModel):
"""Books the library holds that all look like copies of one book."""
books: list[DuplicateBookRead]
class DuplicateDismissal(BaseModel):
"""Two books a reader is saying are not the same book."""
book_a_id: int
book_b_id: int
class BooksUploadResult(BaseModel):
"""The outcome of a multi-file upload: what was created, and what was skipped."""
created: list["BookRead"]
skipped: list[DuplicateFileRead]
# Created, not skipped — these are books that went in and look like something the
# library already had. The reader decides what to do about it.
possible_duplicates: list[PossibleDuplicateRead] = Field(default_factory=list)
class BookMetadataUpdate(BaseModel):
title: str | None = None
subtitle: str | None = None
@@ -160,6 +251,20 @@ class BookMetadataUpdate(BaseModel):
return v
class BookMerge(BaseModel):
"""
Fold several books into one.
`metadata` is the reader's resolution of the fields the records disagreed on.
Anything it does not name keeps the survivor's value — merging metadata is a
judgement, so nothing is guessed on the caller's behalf.
"""
survivor_id: int
merged_ids: list[int]
metadata: Optional["BookMetadataUpdate"] = None
class BookProgressCreate(BaseModel):
percentage: float
epub_cfi: str | None = None
@@ -170,11 +275,3 @@ class BookProgressCreate(BaseModel):
device_id: str | None = None
class BookProgressRead(BaseModel):
percentage: float
epub_cfi: str | None = None
epub_xpointer: str | None = None
pdf_page: int | None = None
completed: bool | None = False
device_type: str | None = None
device_id: str | None = None
+36
View File
@@ -0,0 +1,36 @@
from pydantic import BaseModel
class KosyncProgressUpdate(BaseModel):
"""Schema for uploading progress from KOReader."""
document: str
progress: str | None = None
percentage: float
device: str | None = None
device_id: str | None = None
class KosyncProgressRead(BaseModel):
"""Schema for reading progress to KOReader."""
document: str
progress: str | None = None
percentage: float | None = None
device: str | None = None
device_id: str | None = None
class KosyncDeviceRead(BaseModel):
"""Schema for reading device information."""
id: int
user_id: int
api_key: str
name: str
class KosyncDeviceCreate(BaseModel):
"""Schema for creating a new device."""
name: str
+56 -1
View File
@@ -1,6 +1,7 @@
from pathlib import Path
from typing import Annotated
from pydantic import BaseModel, Field, computed_field
from pydantic import BaseModel, ConfigDict, Field, SkipValidation, computed_field
from litestar.datastructures import UploadFile
from advanced_alchemy.utils.text import slugify
class LibraryCreate(BaseModel):
@@ -8,6 +9,7 @@ class LibraryCreate(BaseModel):
root_path: str
path_template: str | None = "{author}/{title}"
description: str | None = None
icon: str = "library"
read_only: bool = False
@computed_field
@@ -21,13 +23,66 @@ class LibraryRead(BaseModel):
root_path: str
path_template: str
description: str | None
icon: str
read_only: bool
total: int | None = None
class CalibreArchiveUpload(BaseModel):
"""
A zipped Calibre library.
The only way in through the API: a desktop Calibre install is usually not on the
server at all. Importing from a path the server can already see is a server-side
operation, and stays one — `litestar calibre-import` does that.
"""
archive: Annotated[UploadFile, SkipValidation]
allow_duplicates: bool = False
model_config = ConfigDict(arbitrary_types_allowed=True)
class ImportFailureRead(BaseModel):
"""A book the import could not store."""
model_config = ConfigDict(from_attributes=True)
calibre_id: int
title: str
reason: str
class CalibreImportRead(BaseModel):
"""A running or finished import."""
model_config = ConfigDict(from_attributes=True)
id: str
library_id: int
source: str
state: str
total: int
processed: int
created: int
skipped: int
failed: int
current_title: str | None = None
failures: list[ImportFailureRead] = Field(default_factory=list)
# A count, not the records. The library's duplicates screen is what shows them.
possible_duplicates: int = 0
# Set when the run itself broke, as opposed to individual books failing.
error: str | None = None
class LibraryUpdate(BaseModel):
name: str | None
root_path: str | None
path_template: str | None
description: str | None
icon: str | None
read_only: bool | None
+2
View File
@@ -6,3 +6,5 @@ from .author import AuthorService
from .tag import TagService
from .publisher import PublisherService
from .book_progress import BookProgressService
from .kosync_device import KosyncDeviceService
from .kosync_progress import KosyncProgressService
File diff suppressed because it is too large Load Diff
+570
View File
@@ -0,0 +1,570 @@
# src/chitai/services/calibre.py
"""
Read a Calibre library.
This knows about `metadata.db` and the tree beside it, and nothing about `Book`,
`BookService` or a database session — it is a file-format reader, and it is testable
without Postgres or an app. Interpretation belongs to whoever imports what it returns:
identifiers come back exactly as Calibre wrote them, not folded onto Chitai's schemes.
Things about Calibre that are load-bearing here:
- **Never query the views.** `meta` and the `tag_browser_*` family call SQLite functions
Calibre registers from Python at connection time, so `SELECT * FROM meta` fails with
`no such function: sortconcat`. Only base tables are touched below.
- **An unknown date is a sentinel, not a null** — `0101-01-01`, Calibre's
`UNDEFINED_DATE`. It parses fine as a date, so nothing complains; it just makes every
book without a publication date look like it was published in the year 101.
- **`data.name` is not the title.** It is the on-disk stem, truncated to Calibre's
filename limit and sanitised, so the file is `The Project Gutenberg eBook #33283_
Calcul - Silvanus Phillips Thompson.pdf` for a book titled `The Project Gutenberg
eBook #33283: Calculus Made Easy, 2nd Edition`. Names locate files; the database
carries the metadata.
- **`authors.name` escapes a comma as `|`**, which Calibre reverses on read
(`AuthorsTable.unserialize` in its `db/tables.py`).
- **`series_index` defaults to 1.0 whether or not the book is in a series**, so a
position is only meaningful alongside a series.
- **`books_pages_link` is recent and often empty.** It is treated as optional both ways:
the table may not exist, and where it does the rows are frequently `pages = 0` with
`needs_scan = 1`.
Nothing walks the tree: every file is located through `books.path`, which is why
`.caltrash` — where Calibre keeps deleted books, still on disk — cannot be picked up by
accident.
"""
from __future__ import annotations
import asyncio
import shutil
import sqlite3
import tempfile
import zipfile
from collections import defaultdict
from collections.abc import Callable
from dataclasses import dataclass, field
from datetime import date, datetime
from html.parser import HTMLParser
from pathlib import Path, PurePosixPath
METADATA_DB = "metadata.db"
COVER_FILENAME = "cover.jpg"
# Calibre writes `0101-01-01` for "no date". Any year this early is that sentinel rather
# than a publication date somebody meant.
EARLIEST_REAL_YEAR = 1000
# The sidecars a WAL-mode database keeps beside itself. Copied along with it so the
# snapshot can be recovered, since Calibre may be running while this reads.
_DATABASE_SIDECARS = ("-wal", "-shm")
class CalibreLibraryError(Exception):
"""The library cannot be read at all — wrong directory, or no catalogue in it."""
# How far down an archive to look for `metadata.db`. Zipping a Calibre library gives
# either the directory itself or its contents, and a file manager may add a wrapper
# folder on top, so two levels of nesting is normal and more is somebody's backup tree.
_ARCHIVE_SEARCH_DEPTH = 3
# Extraction is refused unless the destination has the uncompressed size plus this
# much headroom. Filling the disk would take the whole application down, not just the
# import.
_DISK_HEADROOM = 256 * 1024 * 1024 # 256 MiB
def _archive_members(archive: zipfile.ZipFile) -> list[zipfile.ZipInfo]:
"""
The entries worth extracting, refusing any that would escape the destination.
`ZipFile.extract` does sanitise names, but relying on that silently is how the next
person to swap the extraction call reintroduces zip-slip. An archive naming
`../../etc/anything` is malformed or hostile, and either way there is nothing to
salvage by continuing.
Raises:
CalibreLibraryError: If any entry points outside the archive root.
"""
members = []
for member in archive.infolist():
if member.is_dir():
continue
name = PurePosixPath(member.filename)
if name.is_absolute() or ".." in name.parts:
raise CalibreLibraryError(
f"The archive contains an entry outside itself: {member.filename!r}"
)
members.append(member)
return members
async def extract_calibre_archive(archive: Path, destination: Path) -> Path:
"""
Unpack a zipped Calibre library and find the catalogue inside it.
Args:
archive: The `.zip` to unpack.
destination: An empty directory to unpack into. The caller owns it and is
responsible for removing it.
Returns:
The directory holding `metadata.db`, which is what `CalibreLibrary` takes.
Raises:
CalibreLibraryError: If the file is not a zip, names an entry outside itself,
would not fit on disk, or holds no `metadata.db`.
"""
return await asyncio.to_thread(_extract_calibre_archive, archive, destination)
def _extract_calibre_archive(archive: Path, destination: Path) -> Path:
if not zipfile.is_zipfile(archive):
raise CalibreLibraryError(
"That is not a zip file. A Calibre library has to be zipped, not tarred."
)
with zipfile.ZipFile(archive) as opened:
members = _archive_members(opened)
if not any(
PurePosixPath(member.filename).name == METADATA_DB for member in members
):
raise CalibreLibraryError(
f"The archive holds no {METADATA_DB}, so it is not a Calibre library"
)
# Checked before writing rather than discovered part-way through: a full disk
# takes the whole application down, and the number is in the archive already.
needed = sum(member.file_size for member in members)
free = shutil.disk_usage(destination).free
if needed + _DISK_HEADROOM > free:
raise CalibreLibraryError(
f"Unpacking needs {needed // (1024 * 1024)} MiB and only "
f"{free // (1024 * 1024)} MiB is free"
)
opened.extractall(destination, members=members)
return _find_catalogue(destination)
def _find_catalogue(root: Path) -> Path:
"""The shallowest directory under `root` holding a `metadata.db`."""
candidates = sorted(
(path.parent for path in root.rglob(METADATA_DB) if path.is_file()),
key=lambda path: len(path.relative_to(root).parts),
)
for candidate in candidates:
if len(candidate.relative_to(root).parts) <= _ARCHIVE_SEARCH_DEPTH:
return candidate
raise CalibreLibraryError(
f"No {METADATA_DB} within {_ARCHIVE_SEARCH_DEPTH} levels of the archive root"
)
@dataclass(frozen=True)
class CalibreFile:
"""One row of Calibre's `data` table: a book in one format."""
path: Path
"""Absolute path, resolved against the library root. Not checked for existence."""
format: str
"""As Calibre stores it, upper case: `EPUB`, `PDF`, `AZW3`."""
size: int
"""`data.uncompressed_size` — Calibre's claim, not a fresh stat."""
@dataclass(frozen=True)
class CalibreBook:
"""One book, with everything Chitai has a column for and nothing it does not."""
calibre_id: int
uuid: str
title: str
authors: list[str] = field(default_factory=list)
description: str | None = None
published_date: date | None = None
series: str | None = None
series_position: str | None = None
tags: list[str] = field(default_factory=list)
publisher: str | None = None
language: str | None = None
identifiers: dict[str, str] = field(default_factory=dict)
"""Keyed by `identifiers.type` verbatim — `isbn`, `mobi-asin`, `amazon`."""
pages: int | None = None
cover: Path | None = None
files: list[CalibreFile] = field(default_factory=list)
class CalibreLibrary:
"""
A Calibre library on disk, opened for reading.
The catalogue is **copied** before it is read. Calibre may be running and writing,
and opening the live file either sees a torn state or needs to recover a write-ahead
log, which read-only access cannot do. The copy is small — hundreds of kilobytes for
a handful of books, single-digit megabytes for thousands — so this costs nothing and
removes the question. The original is never opened by SQLite at all.
"""
def __init__(self, root: Path | str) -> None:
self.root = Path(root)
self._connection: sqlite3.Connection | None = None
self._workspace: Path | None = None
# Every query runs in a worker thread, and `asyncio.to_thread` hands out
# whichever one is free — so the connection outlives the thread that opened it
# and `check_same_thread` has to be off. The lock is what makes that safe: it
# keeps two queries off the connection at once, which is the thing that check
# was standing in for.
self._lock = asyncio.Lock()
@property
def database(self) -> Path:
return self.root / METADATA_DB
async def open(self) -> None:
"""
Copy the catalogue aside and connect to the copy.
Raises:
CalibreLibraryError: If there is no `metadata.db` under the root.
"""
if self._connection is not None:
return
if not await asyncio.to_thread(self.database.is_file):
raise CalibreLibraryError(
f"No {METADATA_DB} in '{self.root}' — that is not a Calibre library"
)
self._workspace = Path(await asyncio.to_thread(tempfile.mkdtemp))
copy = self._workspace / METADATA_DB
await asyncio.to_thread(self._copy_database, copy)
# Read-write on our own copy, deliberately: that is what lets SQLite recover a
# write-ahead log the source may have been mid-way through.
self._connection = sqlite3.connect(str(copy), check_same_thread=False)
def _copy_database(self, destination: Path) -> None:
shutil.copy2(self.database, destination)
for suffix in _DATABASE_SIDECARS:
sidecar = self.database.with_name(self.database.name + suffix)
if sidecar.is_file():
shutil.copy2(sidecar, destination.with_name(destination.name + suffix))
async def close(self) -> None:
"""
Disconnect and remove the copy. Safe to call more than once.
The copy is removed even if closing the connection fails — otherwise a failure
here leaves a catalogue-sized file in the temp directory, and the caller that
failed is exactly the one that will not come back to tidy up.
"""
try:
if self._connection is not None:
async with self._lock:
self._connection.close()
self._connection = None
finally:
if self._workspace is not None:
await asyncio.to_thread(shutil.rmtree, self._workspace, True)
self._workspace = None
async def __aenter__(self) -> CalibreLibrary:
await self.open()
return self
async def __aexit__(self, *_exception: object) -> None:
await self.close()
async def count(self) -> int:
"""How many books the catalogue holds, without reading any of them."""
rows = await self._in_thread(lambda: self._execute("SELECT count(*) FROM books"))
return int(rows[0][0])
async def books(self) -> list[CalibreBook]:
"""
Read the whole catalogue.
One query per table and the joining done in Python, rather than a per-book query
across ten tables. Everything Chitai stores about a book is small, so a
self-hosted catalogue fits in memory comfortably.
Returns:
Every book, in Calibre id order.
"""
return await self._in_thread(self._read_books)
async def _in_thread[T](self, work: Callable[[], T]) -> T:
"""Run one unit of SQLite work off the event loop, and only one at a time."""
async with self._lock:
return await asyncio.to_thread(work)
def _execute(self, statement: str) -> list[tuple]:
if self._connection is None:
raise CalibreLibraryError("The library is not open")
return self._connection.execute(statement).fetchall()
def _has_table(self, name: str) -> bool:
return bool(
self._execute(
f"SELECT 1 FROM sqlite_master WHERE type = 'table' AND name = '{name}'"
)
)
def _read_books(self) -> list[CalibreBook]:
authors = self._grouped(
"SELECT bal.book, a.name FROM books_authors_link bal "
"JOIN authors a ON a.id = bal.author ORDER BY bal.id"
)
tags = self._grouped(
"SELECT btl.book, t.name FROM books_tags_link btl "
"JOIN tags t ON t.id = btl.tag ORDER BY t.name"
)
# Calibre's link tables are unique per book for these two, so the last write
# wins and there is nothing to choose between.
series = self._mapped(
"SELECT bsl.book, s.name FROM books_series_link bsl "
"JOIN series s ON s.id = bsl.series"
)
publishers = self._mapped(
"SELECT bpl.book, p.name FROM books_publishers_link bpl "
"JOIN publishers p ON p.id = bpl.publisher"
)
# A book can carry several languages; Chitai holds one, so the first wins.
languages = self._grouped(
"SELECT bll.book, l.lang_code FROM books_languages_link bll "
"JOIN languages l ON l.id = bll.lang_code ORDER BY bll.item_order"
)
descriptions = self._mapped("SELECT book, text FROM comments")
identifiers: dict[int, dict[str, str]] = defaultdict(dict)
for book_id, name, value in self._execute(
"SELECT book, type, val FROM identifiers"
):
if name and value:
identifiers[book_id][str(name)] = str(value)
files: dict[int, list[tuple[str, str, int]]] = defaultdict(list)
for book_id, format, name, size in self._execute(
"SELECT book, format, name, uncompressed_size FROM data ORDER BY id"
):
files[book_id].append((str(format), str(name), int(size or 0)))
pages: dict[int, int] = {}
if self._has_table("books_pages_link"):
pages = {
book_id: int(count)
for book_id, count in self._execute(
"SELECT book, pages FROM books_pages_link WHERE pages > 0"
)
}
books = []
for row in self._execute(
"SELECT id, title, pubdate, series_index, path, uuid, has_cover "
"FROM books ORDER BY id"
):
book_id, title, pubdate, series_index, path, uuid, has_cover = row
directory = self.root / Path(str(path))
in_series = series.get(book_id)
books.append(
CalibreBook(
calibre_id=book_id,
uuid=str(uuid or ""),
title=str(title or ""),
authors=[unescape_author(name) for name in authors.get(book_id, [])],
description=strip_html(descriptions.get(book_id)),
published_date=parse_date(pubdate),
series=in_series,
# Meaningless without a series: Calibre defaults the index to 1.0 for
# every book, in a series or not.
series_position=(
format_series_index(series_index) if in_series else None
),
tags=tags.get(book_id, []),
publisher=publishers.get(book_id),
language=next(iter(languages.get(book_id, [])), None),
identifiers=dict(identifiers.get(book_id, {})),
pages=pages.get(book_id),
cover=directory / COVER_FILENAME if has_cover else None,
files=[
CalibreFile(
path=directory / f"{name}.{format.lower()}",
format=format,
size=size,
)
for format, name, size in files.get(book_id, [])
],
)
)
return books
def _grouped(self, statement: str) -> dict[int, list[str]]:
"""Run a `(book, value)` query into one list per book, keeping row order."""
grouped: dict[int, list[str]] = defaultdict(list)
for book_id, value in self._execute(statement):
if value is not None:
grouped[book_id].append(str(value))
return grouped
def _mapped(self, statement: str) -> dict[int, str]:
"""Run a `(book, value)` query into one value per book."""
return {
book_id: str(value)
for book_id, value in self._execute(statement)
if value is not None
}
def unescape_author(name: str) -> str:
"""
Undo Calibre's comma escaping.
`authors.name` stores a comma as `|`, and Calibre reverses it on the way out. Left
alone, `Doyle, Sir Arthur Conan` comes back as `Doyle| Sir Arthur Conan`.
"""
return name.replace("|", ",").strip()
def parse_date(value: object) -> date | None:
"""
Read one of Calibre's timestamps, discarding its "unknown" sentinel.
Args:
value: The stored column, which is text in practice but need not be.
Returns:
The date, or None for a null, an unparseable value, or Calibre's
`0101-01-01` placeholder.
"""
if value is None:
return None
if isinstance(value, datetime):
parsed = value.date()
elif isinstance(value, date):
parsed = value
else:
text = str(value).strip()
if not text:
return None
try:
parsed = datetime.fromisoformat(text).date()
except ValueError:
try:
parsed = date.fromisoformat(text[:10])
except ValueError:
return None
return parsed if parsed.year >= EARLIEST_REAL_YEAR else None
def format_series_index(index: object) -> str | None:
"""
Render `series_index` as the string `Book.series_position` holds.
Calibre stores a REAL, so volume seven arrives as `7.0` — which would be stored
verbatim and then compared as a string against the `7` everything else writes.
Fractional positions are real and are kept: `1.5` is a novella between two novels.
"""
if index is None:
return None
try:
number = float(index)
except (TypeError, ValueError):
return None
return str(int(number)) if number.is_integer() else f"{number:g}"
class _TextExtractor(HTMLParser):
"""Flatten markup to text, keeping the line breaks that carried meaning."""
# Tags whose boundaries are a line break rather than nothing at all. Without these
# a description of three paragraphs comes out as one run-on sentence.
_BREAKS = frozenset(
{
"p", "br", "div", "li", "tr", "blockquote", "hr",
"h1", "h2", "h3", "h4", "h5", "h6",
}
)
def __init__(self) -> None:
super().__init__(convert_charrefs=True)
self._parts: list[str] = []
def handle_data(self, data: str) -> None:
self._parts.append(data)
def handle_starttag(self, tag: str, _attrs: object) -> None:
self._break(tag)
def handle_endtag(self, tag: str) -> None:
self._break(tag)
def _break(self, tag: str) -> None:
"""
End the current line, once.
Both halves of `</p><p>` are a boundary, and the open tag of the very first
block is not one at all — so emitting a newline per tag turns two paragraphs
into two blank-line-separated ones with a leading gap. One break per boundary
is what the plain text wants.
"""
if tag in self._BREAKS and self._parts and not self._parts[-1].endswith("\n"):
self._parts.append("\n")
@property
def text(self) -> str:
lines = [line.strip() for line in "".join(self._parts).splitlines()]
return "\n".join(line for line in lines if line).strip()
def strip_html(html: str | None) -> str | None:
"""
Turn Calibre's `comments` into plain text.
`comments.text` is always HTML, and `Book.description` is rendered as text — so the
tags would show literally on the book page.
Args:
html: The stored comment, if there is one.
Returns:
The text, or None when there was nothing or nothing survived.
"""
if not html:
return None
parser = _TextExtractor()
parser.feed(html)
parser.close()
return parser.text or None
@@ -0,0 +1,239 @@
# src/chitai/services/calibre_import.py
"""
Run a Calibre import in the background and report on it.
The import outlives its request — a real library takes minutes to hours — so the handler
starts a task and hands back a handle to poll. The work itself is
`BookService.create_many_from_calibre`; everything here is lifecycle: state, progress,
cancellation, and a session of its own.
**This registry lives in memory, and therefore assumes one worker process.** That holds
today: the production `CMD` is `litestar run`, which is single-process, and the consume
watcher is already an in-process singleton with the same constraint. `TODO.md` records
that the production image should move to uvicorn with a worker count — the day that
happens, a poll can land on a worker that has never heard of the job, and this needs an
`import_jobs` table instead. It is written down here because nothing else will say so.
"""
from __future__ import annotations
import asyncio
import shutil
import uuid
from dataclasses import dataclass, field
from enum import StrEnum
from pathlib import Path
from chitai.config import settings
from chitai.database.models import Library
from chitai.services.book import (
BookService,
CalibreImportProgress,
CalibreImportResult,
UnimportedBook,
)
from chitai.services.calibre import CalibreLibrary
class ImportState(StrEnum):
"""Where a job has got to."""
RUNNING = "running"
FINISHED = "finished"
# Stopped on request. What it imported is complete.
CANCELLED = "cancelled"
# The run itself broke — an unreadable catalogue, a missing library. Distinct from
# individual books failing, which `failures` carries and which never stops the run.
FAILED = "failed"
@dataclass
class ImportJob:
"""One import, running or finished."""
id: str
library_id: int
source: str
state: ImportState = ImportState.RUNNING
total: int = 0
processed: int = 0
created: int = 0
skipped: int = 0
failed: int = 0
current_title: str | None = None
failures: list[UnimportedBook] = field(default_factory=list)
# Books imported that look like something the library already had. A count, not the
# records: the duplicates screen is what shows them, and a big import would make
# this the largest thing in the response for no benefit.
possible_duplicates: int = 0
# Why the whole run stopped, when `state` is FAILED.
error: str | None = None
# The directory the job owns and must delete when it ends: what the uploaded archive
# was unpacked into.
workspace: Path | None = None
_stop: bool = False
@property
def finished(self) -> bool:
return self.state is not ImportState.RUNNING
def absorb(self, result: CalibreImportResult) -> None:
"""Take the final counts from a finished run."""
self.total = result.total
self.created = len(result.created)
self.skipped = len(result.skipped)
self.failed = len(result.failed)
self.failures = list(result.failed)
self.possible_duplicates = len(result.possible_duplicates)
self.current_title = None
self.state = ImportState.CANCELLED if result.stopped else ImportState.FINISHED
class CalibreImportRegistry:
"""
The imports this process knows about.
One instance, held at module scope below. Jobs are kept after they finish so the
screen that started one can still read its result; nothing evicts them, which is
fine for a handful of one-time migrations and is the other reason a table would be
the answer if this ever needed to be durable.
"""
def __init__(self) -> None:
self._jobs: dict[str, ImportJob] = {}
# Held only to keep the tasks referenced. Without this the event loop is free to
# garbage-collect a running task mid-import.
self._tasks: set[asyncio.Task] = set()
def get(self, job_id: str) -> ImportJob | None:
return self._jobs.get(job_id)
def running_for(self, library_id: int) -> ImportJob | None:
"""The unfinished import for a library, if it has one."""
return next(
(
job
for job in self._jobs.values()
if job.library_id == library_id and not job.finished
),
None,
)
def cancel(self, job_id: str) -> ImportJob | None:
"""
Ask a job to stop after the book it is on.
Not `task.cancel()`: that would abandon a book mid-copy, leaving files on disk
with no row describing them. The flag is read between books.
"""
job = self._jobs.get(job_id)
if job is not None and not job.finished:
job._stop = True
return job
def start(
self,
library: Library,
source: Path,
workspace: Path,
label: str,
allow_duplicates: bool = False,
) -> ImportJob:
"""
Begin importing, and return the handle to poll.
Args:
library: The library to import into.
source: The unpacked Calibre library's directory.
workspace: A directory the job owns and deletes when it ends — what the
uploaded archive was unpacked into. The books have been copied into the
library by then, so nothing is lost with it.
label: What to report as the source. `source` is a temp directory that would
mean nothing to the reader, so this is the archive's own name.
allow_duplicates: Import books whose files are already stored.
Returns:
The job, already running.
"""
job = ImportJob(
id=str(uuid.uuid4()),
library_id=library.id,
source=label,
workspace=workspace,
)
self._jobs[job.id] = job
task = asyncio.create_task(self._run(job, library.id, source, allow_duplicates))
self._tasks.add(task)
task.add_done_callback(self._tasks.discard)
return job
async def _run(
self, job: ImportJob, library_id: int, source: Path, allow_duplicates: bool
) -> None:
"""
Do the import, recording everything on the job.
Opens a **session of its own**: the request that started this is long gone, and
its session was closed with it.
"""
from chitai.services.library import LibraryService
def on_progress(progress: CalibreImportProgress) -> None:
job.total = progress.total
job.processed = progress.processed
job.current_title = progress.title
if progress.outcome == "created":
job.created += 1
elif progress.outcome == "skipped":
job.skipped += 1
else:
job.failed += 1
catalogue = CalibreLibrary(source)
try:
await catalogue.open()
async with settings.alchemy_config.get_session() as session:
library = await LibraryService(session=session).get(library_id)
result = await BookService(session=session).create_many_from_calibre(
catalogue,
library,
allow_duplicates=allow_duplicates,
on_progress=on_progress,
should_stop=lambda: job._stop,
)
job.absorb(result)
except Exception as exc:
job.state = ImportState.FAILED
job.error = f"{type(exc).__name__}: {exc}"
finally:
await catalogue.close()
# An unpacked archive is a second copy of the whole library, and the books
# worth keeping have been copied into the library proper by now. Removed
# even when the run failed — especially then, since nothing will come back
# for it.
if job.workspace is not None:
await asyncio.to_thread(shutil.rmtree, job.workspace, True)
registry = CalibreImportRegistry()
+22 -2
View File
@@ -1,6 +1,7 @@
import asyncio
from pathlib import Path
from collections import defaultdict
from chitai.config import settings
from chitai.database.models.library import Library
from chitai.services import BookService, LibraryService
from chitai.services.metadata_extractor import Extractor
@@ -111,13 +112,32 @@ class ConsumeDirectoryWatcher:
"""Process a batch of files."""
try:
books = await self.book_service.create_many_from_existing_files(
result = await self.book_service.create_many_from_existing_files(
list(file_paths),
self.watch_path / Path(library_slug),
library=await self._get_library(library_slug),
)
print(f"Created {len(books)} books!")
print(f"Created {len(result.books)} books!")
if result.duplicates:
print(
f"Moved {len(result.duplicates)} already-stored file(s) "
f"to {settings.duplicate_path}"
)
# Imported all the same — a metadata match is a guess, and there is nobody
# here to ask. The library's duplicates screen is where these get decided.
for possible in result.possible_duplicates:
names = ", ".join(
f"{candidate.title} (#{candidate.book_id})"
for candidate in possible.candidates
)
print(
f"Imported {possible.title!r} (#{possible.book_id}), which may "
f"already be in the library as: {names}"
)
except Exception as e:
print(f"Error processing batch: {e}")
+65 -23
View File
@@ -14,7 +14,7 @@ from advanced_alchemy.extensions.litestar.providers import (
from advanced_alchemy.exceptions import NotFoundError
from advanced_alchemy.filters import CollectionFilter, StatementFilter
from advanced_alchemy.service import FilterTypeT
from sqlalchemy.orm import selectinload, with_loader_criteria
from sqlalchemy.orm import selectinload
from sqlalchemy.ext.asyncio import AsyncSession
from litestar import Request
from litestar.params import Dependency, Parameter
@@ -36,6 +36,8 @@ from chitai.services import (
TagService,
AuthorService,
PublisherService,
KosyncDeviceService,
KosyncProgressService,
)
from chitai.config import settings
from chitai.services.filters.book import (
@@ -49,8 +51,19 @@ from chitai.services.filters.book import (
async def provide_book_service(
db_session: AsyncSession, current_user: m.User | None = None
db_session: AsyncSession, current_user: m.User = Dependency(skip_validation=True)
) -> AsyncGenerator[BookService, None]:
"""
Provide a BookService with per-user data scoped to the caller.
`current_user` is a required dependency, not an optional argument. It used to
default to None with the scoping below wrapped in `if current_user:` — and
when it was not injected, that block was silently skipped, so
`Book.progress_records` and `Book.list_links` loaded *every* user's rows.
`Book.progress` returns `progress_records[0]`, so one user could see another
user's reading position; the shelf checkboxes leaked the same way. Failing
loudly on a missing user is the point of the change.
"""
load = [
selectinload(m.Book.author_links).selectinload(m.BookAuthorLink.author),
selectinload(m.Book.tag_links).selectinload(m.BookTagLink.tag),
@@ -58,30 +71,25 @@ async def provide_book_service(
m.Book.files,
m.Book.identifiers,
m.Book.series,
]
# Load in specific user-book data
if current_user:
# Load progress data
load.extend(
[
selectinload(m.Book.progress_records),
with_loader_criteria(
m.BookProgress, m.BookProgress.user_id == current_user.id
# Reading progress, restricted to the caller.
#
# The restriction lives on the relationship via .and_() rather than in a
# separate with_loader_criteria(). advanced_alchemy's
# get_abstract_loader_options() keeps only _AbstractLoad,
# InstrumentedAttribute, RelationshipProperty and "*" entries and drops
# everything else — and with_loader_criteria() is none of those, so the
# previous criteria were discarded before reaching a query. A
# selectinload() carrying its own .and_() survives that filter.
selectinload(
m.Book.progress_records.and_(m.BookProgress.user_id == current_user.id)
),
]
# Bookshelf membership, restricted to the caller
selectinload(
m.Book.list_links.and_(
m.BookListLink.book_list.has(m.BookList.user_id == current_user.id)
)
# Load shelf data
load.extend(
[
selectinload(m.Book.list_links).selectinload(m.BookListLink.book_list),
with_loader_criteria(
m.BookListLink,
m.Book.lists.any(m.BookList.user_id == current_user.id),
),
).selectinload(m.BookListLink.book_list),
]
)
provider_func = create_service_provider(
BookService,
@@ -119,6 +127,31 @@ def create_book_filter_dependencies(
# Get base filters first
filters = create_filter_dependencies(config, dep_defaults)
# OVERRIDE: id filter typed by the configured id type, not always `str`
#
# advanced_alchemy's `provide_id_filter` annotates `ids` as `list[str]` and
# ignores `config["id_filter"]` entirely, so `?ids=12` reaches the database as
# the string "12" and Postgres refuses `bigint = character varying`. Nothing
# called `?ids=` until the duplicates screen needed to fetch a handful of books
# by id, which is why it went unnoticed.
if id_type := config.get("id_filter"):
id_field = config.get("id_field", "id")
def provide_typed_id_filter(
ids=Parameter(query="ids", default=None, required=False),
) -> CollectionFilter:
return CollectionFilter(field_name=id_field, values=ids)
# Attached as a type object rather than written as an annotation: this module
# has `from __future__ import annotations`, so a written one is stored as the
# string "Optional[list[id_type]]" and resolved against module globals, where
# a local named `id_type` does not exist.
provide_typed_id_filter.__annotations__["ids"] = Optional[list[id_type]]
filters[dep_defaults.ID_FILTER_DEPENDENCY_KEY] = Provide(
provide_typed_id_filter, sync_to_thread=False
)
# OVERRIDE: Custom search filter with trigram search
if config.get("search"):
search_fields = config.get("search")
@@ -340,3 +373,12 @@ def provide_optional_user(request: Request[m.User, Token, Any]) -> m.User | None
async def provide_user_via_basic_auth(request: Request[m.User, None, Any]) -> m.User:
return request.user
async def provide_user_via_kosync_auth(request: Request[m.User, None, Any]) -> m.User:
return request.user
provide_kosync_device_service = create_service_provider(KosyncDeviceService)
provide_kosync_progress_service = create_service_provider(KosyncProgressService)
@@ -16,6 +16,43 @@ import chitai.database.models as m
# - Auto-handle missing values (e.g., skip {series}/ if series is empty)
# Characters that cannot survive being interpolated into a path. The forward slash is
# the one that matters: titles legitimately contain it — "AC/DC", "Him/Her" — and the
# template writes the title straight into a directory name, so an unsanitised one
# silently adds a level and puts the book somewhere `book.path` does not describe.
# Calibre strips these from its own on-disk names and keeps the real title in its
# database, which is how an import surfaces them.
_UNSAFE_IN_PATH = re.compile(r"[/\\\x00-\x1f]")
def sanitize_path_component(value: str) -> str:
"""Make one metadata value safe to use as a single directory or file name."""
return _UNSAFE_IN_PATH.sub("_", value).strip()
def _safe_components(book_data: dict) -> dict:
"""
A shallow copy of the metadata with the values a path is built from sanitised.
Only strings are touched, and only the fields the default template interpolates. A
caller's own template can reach anything else in the dict, which is a reason to keep
this conservative rather than to walk the whole structure.
"""
safe = dict(book_data)
for key in ("title", "series", "series_position"):
if isinstance(safe.get(key), str):
safe[key] = sanitize_path_component(safe[key])
if isinstance(safe.get("authors"), list):
safe["authors"] = [
sanitize_path_component(author) if isinstance(author, str) else author
for author in safe["authors"]
]
return safe
default_path_template = """
/{{book.authors[0] if book.authors else 'Unknown'}}
{%- if book.series -%}
@@ -101,7 +138,12 @@ class BookPathGenerator:
"""
result = self.root_path / Path(self.path_template.render(book=book_data))
# Sanitised per value, never on the rendered result: the separators the template
# puts *between* author, series and title are the whole point of it, and only the
# values interpolated into it must not contribute any of their own.
result = self.root_path / Path(
self.path_template.render(book=_safe_components(book_data))
)
# Clean up
result = re.sub(r"/+", "/", str(result)) # Remove consecutive backslashes
+11 -2
View File
@@ -138,7 +138,7 @@ class ProgressFilter(StatementFilter):
m.BookProgress.completed == False,
m.BookProgress.completed.is_(None),
),
m.BookProgress.progress > 0,
m.BookProgress.percentage > 0,
)
)
@@ -154,7 +154,6 @@ class ProgressFilter(StatementFilter):
@dataclass
class FileFilter(StatementFilter):
"""Filter books that are related to the given files."""
file_ids: list[int]
def append_to_statement(
@@ -166,6 +165,16 @@ class FileFilter(StatementFilter):
return super().append_to_statement(statement, model, *args, **kwargs)
@dataclass
class FileHashFilter(StatementFilter):
file_hashes: list[str]
def append_to_statement(self, statement: StatementTypeT, model: type[ModelT], *args, **kwargs) -> StatementTypeT:
statement = statement.where(
m.Book.files.any(m.FileMetadata.hash.in_(self.file_hashes))
)
return super().append_to_statement(statement, model, *args, **kwargs)
@dataclass
class CustomOrderBy(StatementFilter):
@@ -0,0 +1,35 @@
from __future__ import annotations
import secrets
from chitai.database.models.kosync_device import KosyncDevice
from advanced_alchemy.service import SQLAlchemyAsyncRepositoryService, ModelDictT, schema_dump
from advanced_alchemy.repository import SQLAlchemyAsyncRepository
class KosyncDeviceService(SQLAlchemyAsyncRepositoryService[KosyncDevice]):
"""Service for managing KOReader devices."""
API_KEY_LENGTH_IN_BYTES = 8
class Repo(SQLAlchemyAsyncRepository[KosyncDevice]):
""" Repository for KosyncDevice entities."""
model_type = KosyncDevice
repository_type = Repo
async def create(self, data: ModelDictT[KosyncDevice], **kwargs) -> KosyncDevice:
data = schema_dump(data)
data['api_key'] = self._generate_api_key()
return await super().create(data, **kwargs)
async def get_by_api_key(self, api_key: str) -> KosyncDevice:
return await self.get_one(KosyncDevice.api_key == api_key)
async def regenerate_api_key(self, device_id: int) -> KosyncDevice:
device = await self.get(device_id)
api_key = self._generate_api_key()
device.api_key = api_key
return await self.update(device)
def _generate_api_key(self) -> str:
return secrets.token_hex(self.API_KEY_LENGTH_IN_BYTES)
@@ -0,0 +1,51 @@
from __future__ import annotations
from advanced_alchemy.service import SQLAlchemyAsyncRepositoryService
from advanced_alchemy.repository import SQLAlchemyAsyncRepository
from chitai.database.models.kosync_progress import KosyncProgress
class KosyncProgressService(SQLAlchemyAsyncRepositoryService[KosyncProgress]):
"""Service for managing KOReader sync progress."""
class Repo(SQLAlchemyAsyncRepository[KosyncProgress]):
"""Repository for KosyncProgress entities."""
model_type = KosyncProgress
repository_type = Repo
async def get_by_document_hash(self, user_id: int, document: str) -> KosyncProgress | None:
"""Get progress for a specific document and user."""
return await self.get_one_or_none(
KosyncProgress.user_id == user_id,
KosyncProgress.document == document,
)
async def upsert_progress(
self,
user_id: int,
book_id: int,
document: str,
progress: str | None,
percentage: float,
device: str | None = None,
device_id: str | None = None,
) -> KosyncProgress:
"""Create or update progress for a document."""
existing = await self.get_by_document_hash(user_id, document)
data = {
"user_id": user_id,
"book_id": book_id,
"document": document,
"progress": progress,
"percentage": percentage,
"device": device,
"device_id": device_id,
}
if existing:
return await self.update(data, item_id=existing.id)
return await self.create(data)
+207
View File
@@ -0,0 +1,207 @@
# src/chitai/services/matching.py
"""
Normalisation for book-level duplicate detection.
Two copies of one book rarely agree on how it is written down. One says
`The Metamorphosis`, the other `Metamorphosis`; one credits `Kafka, Franz`, the other
`Franz Kafka`; one carries the ISBN-10 and the other the ISBN-13 of the same edition.
These functions reduce each of those to a single key, so the comparison is an equality
test the database can index rather than a similarity score nobody can explain.
Everything here is pure: the keys are computed once and stored on the row (see
`Book.normalized_title`, `Author.normalized_name`, `Identifier.normalized_value`), so
no Postgres extension is needed at query time.
"""
from __future__ import annotations
import re
import unicodedata
from chitai.services.utils import is_valid_isbn, isbn10_to_isbn13
# Asides a title carries that say nothing about which book it is:
# "Frankenstein (Illustrated)", "Dune [Deluxe]".
_BRACKETED = re.compile(r"[(\[{][^)\]}]*[)\]}]")
# Edition and format qualifiers, matched only as a *trailing* run of words. Anchoring
# to the end is what keeps "The Illustrated Man" a book and "Moby Dick Illustrated" a
# format note — a qualifier trails the title, it is never the thing the title is about.
_EDITION_NOISE = re.compile(
r"\s+(?:"
# "2nd edition", but also the compact forms publishers actually print on a
# cover: "2E", "3 Ed", "5e". The number alone is never enough — "Catch 22" is
# a title and must survive.
r"\d+(?:st|nd|rd|th)?\s*(?:edition|edn|ed|e)"
r"|(?:first|second|third|fourth|fifth|sixth|new|revised|expanded|updated|"
r"annotated|illustrated|unabridged|abridged|complete|definitive|deluxe|"
r"anniversary|collectors|international|kindle|paperback|hardcover|hardback|"
r"ebook|audiobook)"
r"(?:\s+(?:and|&)\s+\w+)*"
r"(?:\s+ed(?:ition|n)?)?"
r")$"
)
_LEADING_ARTICLE = re.compile(r"^(?:the|a|an)\s+")
# Anything that is not a letter, a digit or a space, once accents are gone.
_PUNCTUATION = re.compile(r"[^0-9a-z ]+")
_WHITESPACE = re.compile(r"\s+")
# Junk an extractor leaves on the end of a name: the `;` from a `DC:creator` list, the
# `.epub` from a filename the author's name was read out of.
_TRAILING_SEPARATORS = " ;,&/"
_FILE_EXTENSION = re.compile(
r"\.(?:epub|pdf|mobi|azw3?|djvu|fb2|txt|cbz|cbr)$", re.IGNORECASE
)
# `J. R. R.` survives punctuation stripping as three one-letter words; `J.R.R.` as one.
# Joining any run of them makes both `jrr`.
_INITIAL_RUN = re.compile(r"\b(?:[a-z] )+[a-z]\b")
# ISBNs are the same number under several names; everything else keeps its own.
_ISBN_NAMES = {"isbn", "isbn-10", "isbn10", "isbn-13", "isbn13"}
# Generated fresh for every build of a file, so two copies of one book never share one.
# Matching on them would only re-find files the hash check already catches.
_PER_BUILD_NAMES = {"uuid", "urn:uuid"}
# Below this an identifier is not specific enough to be evidence: a Calibre `id` of
# "42" would otherwise pair two unrelated books.
_MIN_IDENTIFIER_LENGTH = 4
def _fold(text: str) -> str:
"""Casefolded, accent-free, punctuation-free, single-spaced."""
decomposed = unicodedata.normalize("NFKD", text.casefold())
unaccented = "".join(c for c in decomposed if not unicodedata.combining(c))
return _WHITESPACE.sub(" ", _PUNCTUATION.sub(" ", unaccented)).strip()
def normalize_title(title: str | None) -> str:
"""
Reduce a title to the key two copies of one book should share.
`Book.subtitle` is already split off by `Extractor.format_book_title`, so only what
is left in the title column is considered here.
Args:
title: The title as it was stored.
Returns:
The comparison key, or an empty string if nothing survives normalisation —
which is the signal not to match on the title at all.
"""
if not title:
return ""
folded = _fold(_BRACKETED.sub(" ", title).replace("&", " and "))
# Repeated because qualifiers stack: "Dune Deluxe Edition Illustrated".
while (trimmed := _EDITION_NOISE.sub("", folded)) != folded:
folded = trimmed
# An article says nothing, but a title that is only an article is not improved by
# having none, and neither is one that noise removal emptied out.
return _LEADING_ARTICLE.sub("", folded, count=1) or folded
def format_author_name(name: str | None) -> str:
"""
Tidy an author's name into the one form the library writes them in.
Distinct from `normalize_author`, which throws away case, accents and spacing to
build a comparison key. This one is what a reader sees, so it keeps everything
that belongs to the name and only removes what an extractor added: a trailing
separator left over from a creator list, a file extension carried in from a
filename, and the `Surname, Given` ordering that EPUBs file names under.
Args:
name: The name as the file or filename gave it.
Returns:
The name to store, or an empty string if there is nothing left of it.
"""
if not name:
return ""
tidied = _FILE_EXTENSION.sub("", name.strip().strip(_TRAILING_SEPARATORS).strip())
if tidied.count(",") == 1:
surname, given = (part.strip() for part in tidied.split(","))
# Only when the part before the comma is a single word. "Dave Thomas, Andy
# Hunt" is two people in one string, and flipping it would invent a third
# person who does not exist. Leaving an unrecognised form alone is the safe
# failure; rewriting it wrongly is not.
if surname and given and " " not in surname:
tidied = f"{given} {surname}"
return _WHITESPACE.sub(" ", tidied).strip()
def normalize_author(name: str | None) -> str:
"""
Reduce an author's name to the key their other books should share.
Deliberately not reduced to surname plus initial: that collides unrelated people,
and a wrong match here is a book pointed at a stranger's shelf.
Args:
name: The name as it was stored, in either `Franz Kafka` or `Kafka, Franz` form.
Returns:
The comparison key, or an empty string if nothing survives normalisation.
"""
if not name:
return ""
# `Kafka, Franz` is one name written backwards. More than one comma is a list, or a
# suffix, and guessing at either does more harm than leaving it alone.
if name.count(",") == 1:
surname, forename = name.split(",")
name = f"{forename.strip()} {surname.strip()}"
folded = _fold(name)
return _INITIAL_RUN.sub(lambda run: run.group().replace(" ", ""), folded)
def normalize_identifier(name: str, value: str) -> str | None:
"""
Reduce one identifier to a `scheme:value` key, if it can carry a match at all.
Args:
name: What kind of identifier it is, as stored.
value: The identifier itself.
Returns:
The key, or None when the identifier is no use for matching: a per-build UUID,
something too short to be evidence, or an ISBN that fails its own checksum.
"""
name = (name or "").strip().casefold()
value = (value or "").strip()
if not name or not value or name in _PER_BUILD_NAMES:
return None
if name in _ISBN_NAMES:
digits = re.sub(r"[^0-9Xx]", "", value).upper()
if not is_valid_isbn(digits):
return None
# One scheme for both forms: a publisher prints whichever it likes, and the
# ISBN-10 and ISBN-13 of an edition are the same number written twice.
isbn = digits if len(digits) == 13 else isbn10_to_isbn13(digits)
return f"isbn:{isbn}" if isbn else None
folded = _fold(value) or value.casefold()
if len(folded) < _MIN_IDENTIFIER_LENGTH:
return None
return f"{name}:{folded}"
+206 -19
View File
@@ -3,7 +3,6 @@
# TODO: Code is a mess. Clean it up and add docstrings
# Standard library
from abc import ABC, abstractmethod
import datetime
from pathlib import Path
from io import BytesIO
@@ -31,6 +30,147 @@ from chitai.services.utils import (
logger = logging.getLogger(__name__)
# Identifier schemes an EPUB can declare, mapped onto the names the rest of the app
# uses. A scheme arrives either as an `opf:scheme` attribute or as a prefix on the
# value itself (`urn:isbn:…`, `calibre:…`), and the two say the same thing.
_IDENTIFIER_SCHEMES = {
"isbn": "isbn",
"isbn10": "isbn-10",
"isbn-10": "isbn-10",
"isbn13": "isbn-13",
"isbn-13": "isbn-13",
"uuid": "uuid",
"calibre": "calibre",
"doi": "doi",
"asin": "asin",
"amazon": "asin",
"mobi-asin": "asin",
"google": "google",
"goodreads": "goodreads",
}
# `scheme:rest`, with an optional `urn:` in front of it.
_SCHEME_PREFIX = re.compile(r"^(?:urn:)?([A-Za-z][A-Za-z0-9.-]*):(.+)$")
_UUID = re.compile(r"^[0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12}$", re.IGNORECASE)
def parse_identifier(value: str, scheme: str | None = None) -> tuple[str, str] | None:
"""
Work out what one raw identifier is, and what it is worth storing as.
EPUBs write the same ISBN as `9780486282114`, `978-0-486-28211-4` and
`urn:isbn:978-0-486-28211-4`, and carry plenty of identifiers that are not ISBNs
at all. Validating the string verbatim keeps only the first form and throws the
rest away, so normalise first and name whatever survives.
Args:
value: The identifier as the file wrote it.
scheme: What the file said it is, if it said anything — an `opf:scheme`
attribute. A prefix on the value takes precedence over this.
Returns:
The `(name, value)` to store, or `None` when there is nothing usable: an
empty value, or one declared to be an ISBN that fails its own checksum.
"""
value = (value or "").strip()
if not value:
return None
name = _IDENTIFIER_SCHEMES.get((scheme or "").strip().casefold())
# An unrecognised prefix is part of the value rather than a scheme —
# "http://example.com/book" is not an identifier called "http".
if (match := _SCHEME_PREFIX.match(value)) and (
prefixed := _IDENTIFIER_SCHEMES.get(match.group(1).casefold())
):
name = prefixed
value = match.group(2).strip()
if name is None or name.startswith("isbn"):
digits = re.sub(r"[^0-9Xx]", "", value).upper()
if is_valid_isbn(digits):
return f"isbn-{len(digits)}", digits
# Something that announced itself as an ISBN and is not one carries no
# information: storing it would link the reader to a page that does not exist.
if name is not None:
return None
return name or ("uuid" if _UUID.match(value) else "id"), value
# Numbered editions, in the forms covers and catalogue records actually use:
# "3rd Edition", "2E", "4e", "8_e", "/6e", "(2nd edition)", "Third International Edition".
_ORDINAL_WORDS = {
"first": 1, "second": 2, "third": 3, "fourth": 4, "fifth": 5, "sixth": 6,
"seventh": 7, "eighth": 8, "ninth": 9, "tenth": 10, "eleventh": 11, "twelfth": 12,
}
# Words that sit between the number and "Edition" and belong to the same statement.
_EDITION_QUALIFIER = (
r"(?:international|global|revised|updated|expanded|anniversary|deluxe|student|"
r"instructors?|annotated|illustrated|reprint)"
)
_EDITION = re.compile(
rf"""
[\s,;:/\-–—(\[]+ # the separator the statement hangs off
(?:
(?P<num>\d{{1,2}})\s*(?:st|nd|rd|th)?[\s_]*
(?:{_EDITION_QUALIFIER}\s+)*(?:edition\b|edn\b|ed\b\.?|e\b)
| (?P<word>{"|".join(_ORDINAL_WORDS)})\s+
(?:{_EDITION_QUALIFIER}\s+)*(?:edition\b|edn\b|ed\b\.?)
)
[\s)\]]* # and its closing bracket, if it had one
""",
re.IGNORECASE | re.VERBOSE,
)
def split_edition(title: str | None) -> tuple[str | None, int | None]:
"""
Separate a numbered edition statement from the title it is written into.
"Fluent Python, 2nd Edition" is one book with a field for the edition, not a
title. Left in place it also splits the library: the second edition never looks
like the first, and neither matches the copy whose file simply did not mention it.
The number is what makes this safe. Nothing is stripped without one, so
"Catch 22" and "Blade Runner 2049" keep their numbers and "Global Edition"
which is a variant, not a numbered edition, and has nowhere to go in an
integer column — is left in the title where it can still be read.
Args:
title: The title as the file or filename gave it.
Returns:
The title without the edition statement, and the edition number. The title
unchanged and None when there is no numbered edition in it, or when removing
it would leave nothing behind.
"""
if not title:
return title, None
if (match := _EDITION.search(title)) is None:
return title, None
edition = (
int(match["num"]) if match["num"] else _ORDINAL_WORDS[match["word"].casefold()]
)
stripped = _EDITION.sub(" ", title)
stripped = re.sub(r"\s{2,}", " ", stripped)
stripped = re.sub(r"\s+([,;:.!?])", r"\1", stripped) # "Works : What" → "Works: What"
stripped = stripped.strip(" ,;:-–—/")
# A title that is only an edition statement is not improved by having none.
if not stripped:
return title, None
return stripped, edition
class FileExtractor(Protocol):
@classmethod
async def extract_metadata(
@@ -54,25 +194,56 @@ class Extractor:
# EPUB tends to give better metadata results over pdf
sorted_files = sorted(files, key=lambda f: Extractor._get_file_priority(f))
# Identifiers accumulate across formats instead of replacing each other. Every
# other field is a single value where the later, better-trusted format simply
# wins, but identifiers are a *collection*: a book holding an EPUB and a PDF
# genuinely carries what both of them declare, and merging the dict wholesale
# threw away everything the earlier format found. An EPUB that declares an
# ASIN, a Google volume id and a Calibre id kept none of them once a PDF
# contributed a single ISBN.
identifiers: dict[str, str] = {}
for file in sorted_files:
match get_file_extension(file):
case "epub":
metadata = metadata | await EpubExtractor.extract_metadata(file)
extracted = await EpubExtractor.extract_metadata(file)
case "pdf":
metadata = metadata | await PdfExtractor.extract_metadata(file)
extracted = await PdfExtractor.extract_metadata(file)
case _:
break
# First writer wins per name, and the files are already ordered by how
# far their metadata can be trusted. A `dc:identifier` the publisher
# declared outranks an ISBN scraped out of a PDF's copyright page, which
# routinely prints the ISBNs of other formats and older editions too.
for name, value in (extracted.pop("identifiers", None) or {}).items():
identifiers.setdefault(name, value)
metadata = metadata | extracted
if identifiers:
metadata["identifiers"] = identifiers
# Get metadata from file names
for file in files:
metadata = FilenameExtractor.extract_metadata(file) | metadata
# Get metadata from filepath
metadata = metadata | FilepathExtractor.extract_metadata(files[0], root_path)
# Get metadata from filepath. Kept on the left so that anything the file
# itself declared outranks a guess made from its directory names — a folder
# called "Fluent Python - Luciano Ramalho" must not overwrite the title the
# EPUB already carries.
metadata = FilepathExtractor.extract_metadata(files[0], root_path) | metadata
# format the title
if metadata.get('title', None):
title, subtitle = Extractor.format_book_title(metadata["title"])
# Before the subtitle split, so the edition cannot be mistaken for one:
# "How Linux Works, 3rd Edition: What Every Superuser Should Know" has to
# lose the edition first for the colon count to mean anything.
title, edition = split_edition(metadata["title"])
if edition is not None:
metadata.setdefault("edition", edition)
title, subtitle = Extractor.format_book_title(title)
metadata["title"] = title
metadata["subtitle"] = subtitle
@@ -155,7 +326,11 @@ class PdfExtractor(FileExtractor):
if isinstance(data, UploadFile):
data = data.file
try:
doc = pypdfium2.PdfDocument(data)
except Exception as e:
logger.error(f"Error extracting metadata from pdf: {e}")
return {}
basic_metadata = doc.get_metadata_dict(skip_empty=False)
metadata["title"] = basic_metadata["Title"]
@@ -229,7 +404,7 @@ class PdfExtractor(FileExtractor):
try:
return datetime.datetime.strptime(date_portion, "%Y%m%d").date()
except Exception as e:
except Exception:
return None
@classmethod
@@ -356,15 +531,23 @@ class EpubExtractor(FileExtractor):
@classmethod
def _extract_identifiers(cls, epub: epub.EpubBook) -> dict[str, str]:
"""
Every `DC:identifier` the file carries, keyed by what kind of thing it is.
Non-ISBN identifiers are kept: `Identifier` is a free-form name/value pair, so
a Calibre id or an ASIN costs nothing to store and is one more thing two copies
of a book can be recognised by.
"""
identifiers = {}
for id in epub.get_metadata("DC", "identifier"):
if is_valid_isbn(id[0]):
if len(id[0]) == 13:
identifiers.update({"isbn-13": id[0]})
for value, attributes in epub.get_metadata("DC", "identifier"):
scheme = None
if isinstance(attributes, dict):
scheme = attributes.get("opf:scheme") or attributes.get("scheme")
elif len(id[0]) == 10:
identifiers.update({"isbn-10": id[0]})
if (parsed := parse_identifier(value, scheme)) is not None:
name, parsed_value = parsed
identifiers[name] = parsed_value
return identifiers
@@ -373,7 +556,7 @@ class EpubExtractor(FileExtractor):
try:
return epub.get_metadata("DC", "description")[0][0]
except:
except Exception:
return None
@classmethod
@@ -382,15 +565,15 @@ class EpubExtractor(FileExtractor):
date_str = epub.get_metadata("DC", "date")[0][0].split("T")[0]
return datetime.date.fromisoformat(date_str)
except:
except Exception:
return None
@classmethod
def _extract_publisher(cls, epub: epub.EpubBook) -> str | None:
try:
epub.get_metadata("DC", "publisher")[0][0]
return epub.get_metadata("DC", "publisher")[0][0]
except:
except Exception:
return None
@classmethod
@@ -414,7 +597,7 @@ class EpubExtractor(FileExtractor):
cover_item = epub.get_item_with_id(cover_id)
if cover_item:
return PIL.Image.open(BytesIO(cover_item.content))
except Exception as e:
except Exception:
pass # Fallback to next strategy
# Strategy 2: Search image filenames for "cover" keyword
@@ -482,7 +665,11 @@ class FilenameExtractor(FileExtractor):
elif isinstance(input, Path):
filename = get_filename(input, ext=False)
elif isinstance(input, UploadFile):
filename = Path(input.filename).name
# `.stem`, not `.name`: the extension is not part of the metadata, and
# this is the browser upload path, so keeping it is how a library fills
# up with authors called "Sam Newman.epub". The other two branches have
# always stripped it.
filename = Path(input.filename).stem
else:
raise ValueError("Input type not supported")
+6 -1
View File
@@ -47,8 +47,13 @@ def convert_book_to_entry(book: m.Book) -> Entry:
link=[
ImageLink(href=f"/{book.cover_image}", type="image/webp"),
*[
# The only place a content type has to be a string: `Link.type` is
# required, and a null fails the whole feed rather than one entry.
# `application/octet-stream` is the registered way to say "opaque
# bytes", which is exactly what an unnamed format is.
AcquisitionLink(
href=f"/opds/download/{book.id}/{file.id}", type=file.content_type
href=f"/opds/download/{book.id}/{file.id}",
type=file.content_type or "application/octet-stream",
)
for file in book.files
],
+280 -1
View File
@@ -1,6 +1,11 @@
# src/chitai/services/utils.py
# Standard library
from __future__ import annotations
import errno
import hashlib
import mimetypes
from pathlib import Path
import shutil
from typing import BinaryIO
@@ -12,6 +17,165 @@ import aiofiles
import aiofiles.os as aios
from litestar.datastructures import UploadFile
##################################
# KOReader file hash utilities #
##################################
# KOReader partial MD5 constants
# These match KOReader's partial MD5 implementation for document identification
# KOReader samples 1024 bytes at specific offsets calculated using 32-bit left shift.
# The shift wrapping behavior (shift & 0x1F) causes i=-1 to produce offset 0.
# Offsets: 0, 1024, 4096, 16384, 65536, 262144, 1048576, ...
KO_STEP = 1024
KO_SAMPLE_SIZE = 1024
KO_INDICES = range(-1, 11) # -1 to 10 inclusive
# How much is read at a time while hashing.
HASH_CHUNK_SIZE = 262144 # 256 KiB
def _lshift32(val: int, shift: int) -> int:
"""
32-bit left shift matching LuaJIT's bit.lshift behavior.
LuaJIT masks the shift amount to 5 bits (0-31) and performs 32-bit arithmetic.
This causes negative shifts to wrap: shift=-2 becomes shift=30, and
1024 << 30 overflows 32 bits to produce 0.
"""
val &= 0xFFFFFFFF
shift &= 0x1F
return (val << shift) & 0xFFFFFFFF
def _get_koreader_offsets() -> list[int]:
"""Get all KOReader sampling offsets."""
return [_lshift32(KO_STEP, 2 * i) for i in KO_INDICES]
def _partial_md5_from_chunk(
chunk: bytes,
hasher: hashlib._Hash,
offsets: list[int],
chunk_start: int,
) -> None:
"""
Update partial MD5 hasher with sampled bytes from a chunk.
KOReader samples 1024 bytes at specific offsets rather than hashing
the entire file. This function checks if any sampling offsets fall
within the current chunk and updates the hasher with those bytes.
Args:
chunk: The current chunk of file data.
hasher: The MD5 hasher to update.
offsets: List of byte offsets to sample from the file.
chunk_start: The starting byte position of this chunk in the file.
"""
chunk_len = len(chunk)
for offset in offsets:
if chunk_start <= offset < chunk_start + chunk_len:
start = offset - chunk_start
end = min(start + KO_SAMPLE_SIZE, chunk_len)
hasher.update(chunk[start:end])
async def calculate_koreader_hash(file_path: Path) -> str:
"""
Calculate KOReader-compatible partial MD5 hash for a file.
KOReader uses a partial MD5 algorithm that samples 1024 bytes at specific
offsets rather than hashing the entire file. This provides fast document
identification for large ebook files.
The offsets are calculated using 32-bit left shift: 1024 << (2*i) for i from -1 to 10.
Due to 32-bit overflow, i=-1 produces offset 0:
0, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, ...
Args:
file_path: Path to the file to hash.
Returns:
The hexadecimal MD5 hash string.
"""
hasher = hashlib.md5()
offsets = _get_koreader_offsets()
file_pos = 0
async with aiofiles.open(file_path, "rb") as f:
while chunk := await f.read(HASH_CHUNK_SIZE):
_partial_md5_from_chunk(chunk, hasher, offsets, file_pos)
file_pos += len(chunk)
return hasher.hexdigest()
class StreamingHasher:
"""
Helper class for calculating KOReader hash while streaming file data.
Allows hash calculation during file writes without needing to re-read
the file after writing.
"""
def __init__(self) -> None:
self.hasher = hashlib.md5()
self.offsets = _get_koreader_offsets()
self.position = 0
def update(self, chunk: bytes) -> None:
"""Update hash with a chunk of data."""
_partial_md5_from_chunk(chunk, self.hasher, self.offsets, self.position)
self.position += len(chunk)
def hexdigest(self) -> str:
"""Return the final hash."""
return self.hasher.hexdigest()
@property
def size(self) -> int:
"""Total number of bytes fed in so far."""
return self.position
async def fingerprint_upload(file: UploadFile) -> tuple[str, int]:
"""
Calculate the hash and byte size of an uploaded file without storing it.
Duplicate detection has to answer before anything is written to the library, so
the file is read here and rewound for whoever writes it afterwards.
Args:
file: The uploaded file to read.
Returns:
The file's `(hash, size)` pair.
"""
hasher = StreamingHasher()
await file.seek(0)
while chunk := await file.read(HASH_CHUNK_SIZE):
hasher.update(chunk)
await file.seek(0)
return hasher.hexdigest(), hasher.size
async def fingerprint_file(file_path: Path) -> tuple[str, int]:
"""
Calculate the hash and byte size of a file already on disk.
Args:
file_path: Path to the file to read.
Returns:
The file's `(hash, size)` pair.
"""
stats = await aios.stat(file_path)
return await calculate_koreader_hash(file_path), stats.st_size
##################################
# Filesystem related utilities #
##################################
@@ -64,7 +228,39 @@ async def move_file(src_path: Path, dest_path: Path, create_dirs=True) -> None:
if dest_dir: # Only create if there's a directory path
await aios.makedirs(dest_dir, exist_ok=True)
try:
await aios.rename(src_path, dest_path)
except OSError as exc:
if exc.errno != errno.EXDEV:
raise
# Source and destination are on different filesystems, which rename cannot
# cross. Libraries, the consume directory and the duplicates directory are
# all configured separately, so they can easily be separate mounts.
shutil.move(str(src_path), str(dest_path))
async def copy_file(src_path: Path, dest_path: Path, create_dirs: bool = True) -> None:
"""
Copy a file, streaming it rather than reading it whole.
`shutil.copy` would block the event loop for as long as the read takes, which for a
40 MB ebook — and a few thousand of them in a row — is not acceptable.
Args:
src_path: The file to copy. Left exactly as it is.
dest_path: Where the copy goes.
create_dirs: Create the destination's parent directories first.
"""
if create_dirs and dest_path.parent:
await aios.makedirs(dest_path.parent, exist_ok=True)
async with (
aiofiles.open(src_path, "rb") as source,
aiofiles.open(dest_path, "wb") as destination,
):
while chunk := await source.read(HASH_CHUNK_SIZE):
await destination.write(chunk)
async def move_dir_contents(source_dir: Path | str, target_dir: Path | str) -> None:
"""
@@ -287,6 +483,66 @@ def get_filename(file: Path | str, ext: bool = True) -> str:
return filename.stem
# Content types for the ebook formats `mimetypes` does not know. Python's built-in map
# covers `.epub`, `.pdf`, `.azw3`, `.cbz`, `.cbr` and `.djvu`, and answers `None` for
# every format below — so a library imported from elsewhere, which is where MOBI and
# AZW files come from, stores nothing for them.
#
# That matters downstream because an OPDS acquisition link is how a reader app decides
# whether it can open a file at all.
# What a client sends when it does not know either. Treated as an absence rather than
# an answer: storing it would be indistinguishable from having determined a format, and
# it is the value browsers post for every extension they do not recognise.
_UNSPECIFIED = "application/octet-stream"
EBOOK_CONTENT_TYPES = {
"mobi": "application/x-mobipocket-ebook",
"prc": "application/x-mobipocket-ebook",
"azw": "application/vnd.amazon.ebook",
"fb2": "application/x-fictionbook+xml",
"fbz": "application/x-zip-compressed-fb2",
"lit": "application/x-ms-reader",
"lrf": "application/x-sony-bbeb",
"cb7": "application/x-cb7",
}
def guess_content_type(
file: Path | str | UploadFile, fallback: str | None = None
) -> str | None:
"""
Name a file's format from its extension.
The extension is trusted ahead of anything a client said: a browser posts
`application/octet-stream` for every format it does not recognise, which is most
ebook formats, and that answer is worth less than the `.mobi` on the end of the
name.
Args:
file: The file to name, as a path or an upload.
fallback: What to use when neither table knows the extension — a client-supplied
content type, if there is one. `application/octet-stream` is discarded: it
is the client saying it does not know, which is not information.
Returns:
The content type, or None when nothing can name it. Null is the honest answer
and the column is nullable: a caller that structurally needs a string should
substitute one where it needs it, rather than have an invented value stored.
"""
extension = get_file_extension(file)
if known := EBOOK_CONTENT_TYPES.get(extension):
return known
guessed, _ = mimetypes.guess_type(get_filename(file))
if fallback == _UNSPECIFIED:
fallback = None
return guessed or fallback
###############################
# ISBN Validation utilities #
###############################
@@ -312,7 +568,7 @@ def is_valid_isbn(isbn: str) -> bool:
return is_valid_isbn13(isbn)
else:
return False
except:
except Exception:
return False
@@ -337,6 +593,29 @@ def is_valid_isbn10(isbn: str) -> bool:
return str(check_digit) == isbn[-1] or (check_digit == 10 and isbn[-1] in "Xx")
def isbn10_to_isbn13(isbn: str) -> str | None:
"""
Convert an ISBN-10 to the ISBN-13 naming the same edition.
The two are the same number written twice: prefix `978`, drop the ISBN-10 check
digit, recompute the check digit under the ISBN-13 rule. Matching only works if
both forms collapse onto one, since a publisher may print either.
Args:
isbn: A 10-character ISBN, digits and an optional trailing `X` only.
Returns:
The equivalent ISBN-13, or None if the input is not a valid ISBN-10.
"""
if not is_valid_isbn(isbn) or len(isbn) != 10:
return None
digits = f"978{isbn[:9]}"
total = sum(int(digit) * (1 if i % 2 == 0 else 3) for i, digit in enumerate(digits))
return f"{digits}{(10 - total % 10) % 10}"
def is_valid_isbn13(isbn: str) -> bool:
"""
Validate an ISBN-13 number using its check digit.
+282
View File
@@ -0,0 +1,282 @@
"""
Build a Calibre library on disk, for tests to read.
Generated rather than committed as a binary `metadata.db`, because the rows worth
testing are the awkward ones — the year-101 pubdate, a `|` in an author name, a REAL
series index, HTML in a comment — and those are clearer written out in Python than
hidden inside a blob.
The schema below is Calibre's own, copied from a real library's `sqlite_master`, reduced
to the tables the reader touches. `books_pages_link` is created separately by
`add_pages`: it is recent, and a library made by an older Calibre will not have it.
"""
from __future__ import annotations
import shutil
import sqlite3
from pathlib import Path
SCHEMA = """
CREATE TABLE books (
id INTEGER PRIMARY KEY AUTOINCREMENT,
title TEXT NOT NULL DEFAULT 'Unknown' COLLATE NOCASE,
sort TEXT COLLATE NOCASE,
timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
pubdate TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
series_index REAL NOT NULL DEFAULT 1.0,
author_sort TEXT COLLATE NOCASE,
path TEXT NOT NULL DEFAULT '',
uuid TEXT,
has_cover BOOL DEFAULT 0,
last_modified TIMESTAMP NOT NULL DEFAULT '2000-01-01 00:00:00+00:00'
);
CREATE TABLE authors (
id INTEGER PRIMARY KEY, name TEXT NOT NULL COLLATE NOCASE,
sort TEXT COLLATE NOCASE, link TEXT NOT NULL DEFAULT '', UNIQUE(name)
);
CREATE TABLE books_authors_link (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL, author INTEGER NOT NULL,
UNIQUE(book, author)
);
CREATE TABLE publishers (
id INTEGER PRIMARY KEY, name TEXT NOT NULL COLLATE NOCASE,
sort TEXT COLLATE NOCASE, link TEXT NOT NULL DEFAULT '', UNIQUE(name)
);
CREATE TABLE books_publishers_link (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL, publisher INTEGER NOT NULL,
UNIQUE(book)
);
CREATE TABLE tags (
id INTEGER PRIMARY KEY, name TEXT NOT NULL COLLATE NOCASE,
link TEXT NOT NULL DEFAULT '', UNIQUE (name)
);
CREATE TABLE books_tags_link (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL, tag INTEGER NOT NULL,
UNIQUE(book, tag)
);
CREATE TABLE series (
id INTEGER PRIMARY KEY, name TEXT NOT NULL COLLATE NOCASE,
sort TEXT COLLATE NOCASE, link TEXT NOT NULL DEFAULT '', UNIQUE (name)
);
CREATE TABLE books_series_link (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL, series INTEGER NOT NULL,
UNIQUE(book)
);
CREATE TABLE languages (
id INTEGER PRIMARY KEY, lang_code TEXT NOT NULL COLLATE NOCASE,
link TEXT NOT NULL DEFAULT '', UNIQUE(lang_code)
);
CREATE TABLE books_languages_link (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL, lang_code INTEGER NOT NULL,
item_order INTEGER NOT NULL DEFAULT 0, UNIQUE(book, lang_code)
);
CREATE TABLE comments (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL,
text TEXT NOT NULL COLLATE NOCASE, UNIQUE(book)
);
CREATE TABLE identifiers (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL,
type TEXT NOT NULL DEFAULT 'isbn' COLLATE NOCASE,
val TEXT NOT NULL COLLATE NOCASE, UNIQUE(book, type)
);
CREATE TABLE data (
id INTEGER PRIMARY KEY, book INTEGER NOT NULL,
format TEXT NOT NULL COLLATE NOCASE, uncompressed_size INTEGER NOT NULL,
name TEXT NOT NULL, UNIQUE(book, format)
);
"""
# Calibre's own "no date". Stored, never null, and a valid date — which is exactly why
# it has to be recognised rather than parsed.
UNDEFINED_DATE = "0101-01-01 00:00:00+00:00"
# What a `cover.jpg` that PIL cannot read looks like. Real libraries hold these, from
# an interrupted download or a failed conversion.
CORRUPT_COVER = b"\xff\xd8\xff\xe0 not really a jpeg"
def write_cover(path: Path) -> None:
"""
Write a real, readable JPEG.
Generated with PIL rather than embedded as a hex blob: a hand-rolled JPEG that is
subtly malformed fails inside the import as an unrelated error, which is exactly the
confusion this avoids.
"""
from PIL import Image
Image.new("RGB", (2, 3), (10, 20, 30)).save(path, "JPEG")
class CalibreFixture:
"""A Calibre library being assembled under `root`."""
def __init__(self, root: Path) -> None:
self.root = root
self.root.mkdir(parents=True, exist_ok=True)
self.connection = sqlite3.connect(self.root / "metadata.db")
self.connection.executescript(SCHEMA)
def add_pages_table(self) -> None:
"""Add `books_pages_link`, which only a recent Calibre creates."""
self.connection.executescript(
"""
CREATE TABLE books_pages_link (
book INTEGER PRIMARY KEY,
pages INTEGER DEFAULT 0 NOT NULL,
algorithm INTEGER DEFAULT 0 NOT NULL,
format TEXT DEFAULT '' NOT NULL COLLATE NOCASE,
format_size INTEGER DEFAULT 0 NOT NULL,
timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
needs_scan INTEGER NOT NULL DEFAULT 0
);
"""
)
def add_book(
self,
book_id: int,
title: str,
*,
authors: list[str] | None = None,
pubdate: str = UNDEFINED_DATE,
series: str | None = None,
series_index: float = 1.0,
tags: list[str] | None = None,
publisher: str | None = None,
languages: list[str] | None = None,
comment: str | None = None,
identifiers: dict[str, str] | None = None,
uuid: str | None = None,
pages: int | None = None,
cover: bool = False,
corrupt_cover: bool = False,
formats: dict[str, Path] | None = None,
directory: str | None = None,
) -> Path:
"""
Add one book, with its files laid out the way Calibre lays them out.
Args:
formats: Format name (`EPUB`) to a real file to copy in. Its on-disk stem is
Calibre's, not the title — that is the point of the `data` table.
directory: Override the `books.path` value, for testing a row whose
directory is not where the convention would put it.
Returns:
The book's directory.
"""
author_names = authors or ["Unknown"]
relative = directory or f"{author_names[0]}/{title} ({book_id})"
book_directory = self.root / relative
book_directory.mkdir(parents=True, exist_ok=True)
self.connection.execute(
"INSERT INTO books (id, title, pubdate, series_index, path, uuid, has_cover) "
"VALUES (?, ?, ?, ?, ?, ?, ?)",
(
book_id,
title,
pubdate,
series_index,
relative,
uuid or f"uuid-{book_id}",
int(cover or corrupt_cover),
),
)
for name in author_names:
self._link("authors", "books_authors_link", "author", book_id, name)
for name in tags or []:
self._link("tags", "books_tags_link", "tag", book_id, name)
if series:
self._link("series", "books_series_link", "series", book_id, series)
if publisher:
self._link(
"publishers", "books_publishers_link", "publisher", book_id, publisher
)
for order, code in enumerate(languages or []):
language_id = self._lookup("languages", "lang_code", code)
self.connection.execute(
"INSERT INTO books_languages_link (book, lang_code, item_order) "
"VALUES (?, ?, ?)",
(book_id, language_id, order),
)
if comment is not None:
self.connection.execute(
"INSERT INTO comments (book, text) VALUES (?, ?)", (book_id, comment)
)
for name, value in (identifiers or {}).items():
self.connection.execute(
"INSERT INTO identifiers (book, type, val) VALUES (?, ?, ?)",
(book_id, name, value),
)
if pages is not None:
self.connection.execute(
"INSERT INTO books_pages_link (book, pages) VALUES (?, ?)",
(book_id, pages),
)
if corrupt_cover:
(book_directory / "cover.jpg").write_bytes(CORRUPT_COVER)
elif cover:
write_cover(book_directory / "cover.jpg")
for format, origin in (formats or {}).items():
# Calibre's on-disk stem: sanitised, truncated, and not the title.
stem = f"{title[:40]} - {author_names[0]}".replace(":", "_")
destination = book_directory / f"{stem}.{format.lower()}"
shutil.copy(origin, destination)
self.connection.execute(
"INSERT INTO data (book, format, uncompressed_size, name) "
"VALUES (?, ?, ?, ?)",
(book_id, format, destination.stat().st_size, stem),
)
return book_directory
def add_missing_format(self, book_id: int, format: str, stem: str) -> None:
"""Record a file in the catalogue without putting one on disk."""
self.connection.execute(
"INSERT INTO data (book, format, uncompressed_size, name) VALUES (?, ?, ?, ?)",
(book_id, format, 1234, stem),
)
def _link(
self, table: str, link_table: str, column: str, book_id: int, name: str
) -> None:
item_id = self._lookup(table, "name", name)
self.connection.execute(
f"INSERT INTO {link_table} (book, {column}) VALUES (?, ?)",
(book_id, item_id),
)
def _lookup(self, table: str, column: str, value: str) -> int:
row = self.connection.execute(
f"SELECT id FROM {table} WHERE {column} = ?", (value,)
).fetchone()
if row:
return int(row[0])
cursor = self.connection.execute(
f"INSERT INTO {table} ({column}) VALUES (?)", (value,)
)
return int(cursor.lastrowid or 0)
def commit(self) -> Path:
"""Finish writing and return the library root."""
self.connection.commit()
self.connection.close()
return self.root
+27 -1
View File
@@ -40,7 +40,13 @@ pytest_plugins = [
@pytest.fixture(autouse=True)
def _patch_settings(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setattr(settings, "book_cover_path", f"{tmp_path}/covers")
# PIL will not create the directory it is asked to save into, so anything that
# imports a file carrying a cover needs it to exist first.
covers = tmp_path / "covers"
covers.mkdir()
monkeypatch.setattr(settings, "book_cover_path", str(covers))
monkeypatch.setattr(settings, "duplicate_path", str(tmp_path / "duplicates"))
@pytest.fixture(name="engine")
@@ -202,6 +208,26 @@ async def bookshelf_service(
yield service
@pytest.fixture
async def kosync_progress_service(
sessionmaker: async_sessionmaker[AsyncSession],
) -> AsyncGenerator[services.KosyncProgressService, None]:
"""Create KosyncProgressService instance."""
async with sessionmaker() as session:
async with services.KosyncProgressService.new(session) as service:
yield service
@pytest.fixture
async def kosync_device_service(
sessionmaker: async_sessionmaker[AsyncSession],
) -> AsyncGenerator[services.KosyncDeviceService, None]:
"""Create KosyncDeviceService instance."""
async with sessionmaker() as session:
async with services.KosyncDeviceService.new(session) as service:
yield service
# Data fixtures
+497 -2
View File
@@ -37,7 +37,8 @@ from pathlib import Path
(
Path("tests/data_files/Calculus Made Easy - Silvanus Thompson.pdf"),
2,
"The Project Gutenberg eBook #33283: Calculus Made Easy, 2nd Edition",
# The ", 2nd Edition" is split off into `edition`, not kept in the title.
"The Project Gutenberg eBook #33283: Calculus Made Easy",
["Silvanus Phillips Thompson"],
),
],
@@ -208,6 +209,21 @@ async def test_get_book_file(
assert downloaded_content == file_content
async def test_list_books_by_id(populated_authenticated_client: AsyncClient) -> None:
"""
`?ids=` has to reach the database as integers.
advanced_alchemy's stock id filter annotates the parameter as `list[str]` whatever
the configured id type, so the ids arrived as strings and Postgres refused to
compare a bigint primary key against them. Nothing called it until a screen needed
to fetch a handful of books by id.
"""
response = await populated_authenticated_client.get("/books?ids=1&ids=2&pageSize=10")
assert response.status_code == 200
assert sorted(book["id"] for book in response.json()["items"]) == [1, 2]
async def test_get_book_by_id(populated_authenticated_client: AsyncClient) -> None:
"""Test retrieving a specific book by ID."""
@@ -366,7 +382,395 @@ async def test_create_multiple_books_from_directory(
assert response.status_code == 201
data = response.json()
assert len(data.get("items") or data.get("data")) >= 1
assert len(data["created"]) == 2
assert data["skipped"] == []
async def test_create_books_from_parent_directory_keeps_embedded_title(
authenticated_client: AsyncClient,
) -> None:
"""A folder name in the upload path must not override the file's own metadata.
The browser sends webkitRelativePath, so picking the shelf above a book's folder
submits one more path component than picking the folder itself. That extra level
used to make the directory name win over the title inside the EPUB.
"""
source = Path("tests/data_files/Metamorphosis - Franz Kafka.epub")
files = [
(
"files",
(
"Shelf/Metamorphosis - Franz Kafka/Metamorphosis.epub",
source.read_bytes(),
"application/epub+zip",
),
)
]
response = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=files, data={"library_id": 1}
)
assert response.status_code == 201
books = response.json()["created"]
assert len(books) == 1
assert books[0]["title"] == "Metamorphosis"
async def test_create_books_groups_formats_within_one_folder(
authenticated_client: AsyncClient,
) -> None:
"""Picking a book's own folder yields one book with both formats, not two books."""
epub = Path("tests/data_files/Metamorphosis - Franz Kafka.epub").read_bytes()
pdf = Path("tests/data_files/Calculus Made Easy - Silvanus Thompson.pdf").read_bytes()
files = [
("files", ("Metamorphosis/Metamorphosis.epub", epub, "application/epub+zip")),
("files", ("Metamorphosis/Metamorphosis.pdf", pdf, "application/pdf")),
]
response = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=files, data={"library_id": 1}
)
assert response.status_code == 201
books = response.json()["created"]
assert len(books) == 1
assert len(books[0]["files"]) == 2
class TestDuplicateHandling:
"""A file the library already holds must not be stored a second time."""
epub_path = Path("tests/data_files/Metamorphosis - Franz Kafka.epub")
def upload(self, name: str | None = None) -> list[tuple[str, tuple]]:
return [
(
"files",
(
name or self.epub_path.name,
self.epub_path.read_bytes(),
"application/epub+zip",
),
)
]
async def test_bulk_upload_reports_skipped_files(
self, authenticated_client: AsyncClient
) -> None:
"""Re-dropping a folder must import what is new and name what was not."""
first = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload()
)
assert first.status_code == 201
created = first.json()["created"][0]
second = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload()
)
assert second.status_code == 201
result = second.json()
assert result["created"] == []
assert len(result["skipped"]) == 1
skipped = result["skipped"][0]
assert skipped["filename"] == self.epub_path.name
assert skipped["book_id"] == created["id"]
assert skipped["book_title"] == created["title"]
async def test_bulk_upload_can_be_forced(
self, authenticated_client: AsyncClient
) -> None:
await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload()
)
response = await authenticated_client.post(
"/books/fromFiles?library_id=1&allow_duplicates=true", files=self.upload()
)
assert response.status_code == 201
assert len(response.json()["created"]) == 1
assert response.json()["skipped"] == []
async def test_single_book_create_conflicts(
self, authenticated_client: AsyncClient
) -> None:
"""Naming files deliberately earns a refusal rather than a silent drop."""
await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload()
)
response = await authenticated_client.post(
"/books?library_id=1", files=self.upload(), data={"library_id": 1}
)
assert response.status_code == 409
assert response.json()["extra"][0]["filename"] == self.epub_path.name
forced = await authenticated_client.post(
"/books?library_id=1&allow_duplicates=true",
files=self.upload(),
data={"library_id": 1},
)
assert forced.status_code == 201
async def test_adding_another_books_file_conflicts(
self, authenticated_client: AsyncClient
) -> None:
created = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload()
)
book_id = created.json()["created"][0]["id"]
other = await authenticated_client.post(
"/books?library_id=1",
files=[
(
"files",
(
"war.epub",
Path("tests/data_files/The Art of War - Sun Tzu.epub").read_bytes(),
"application/epub+zip",
),
)
],
data={"library_id": 1},
)
other_id = other.json()["id"]
response = await authenticated_client.post(
f"/books/{book_id}/files",
files=[
(
"files",
(
"war.epub",
Path("tests/data_files/The Art of War - Sun Tzu.epub").read_bytes(),
"application/epub+zip",
),
)
],
)
assert response.status_code == 409
assert response.json()["extra"][0]["book_id"] == other_id
async def test_resending_a_books_own_file_changes_nothing(
self, authenticated_client: AsyncClient
) -> None:
created = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload()
)
book_id = created.json()["created"][0]["id"]
response = await authenticated_client.post(
f"/books/{book_id}/files", files=self.upload()
)
assert response.status_code == 201
assert len(response.json()["files"]) == 1
async def test_duplicates_can_be_checked_before_uploading(
self, authenticated_client: AsyncClient
) -> None:
"""The pre-flight check answers from hashes alone, with no file sent."""
created = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload()
)
book = created.json()["created"][0]
stored = book["files"][0]
response = await authenticated_client.post(
"/books/duplicate-files?library_id=1",
json=[
{
"hash": stored["hash"],
"size": stored["size"],
"filename": "local-copy.epub",
},
{"hash": stored["hash"], "size": stored["size"] + 1},
{"hash": "0" * 32, "size": 1234},
],
)
assert response.status_code == 200
matches = response.json()
assert len(matches) == 1
assert matches[0]["filename"] == "local-copy.epub"
assert matches[0]["book_id"] == book["id"]
@pytest.mark.asyncio
class TestDuplicateBooks:
"""A second copy of a book is imported and reported, never refused."""
epub_path = Path("tests/data_files/Metamorphosis - Franz Kafka.epub")
def upload(self, name: str, pad: bool = False) -> list[tuple[str, tuple]]:
"""
The fixture, optionally padded so it is a different file and the same book.
Padding the archive changes its size and its sampled hash without disturbing
the metadata, which is exactly the case the file-level check cannot see.
"""
data = self.epub_path.read_bytes() + (b"\0" * 64 if pad else b"")
return [("files", (name, data, "application/epub+zip"))]
async def test_a_second_edition_is_created_and_reported(
self, authenticated_client: AsyncClient
) -> None:
first = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("first.epub")
)
original = first.json()["created"][0]
second = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("second.epub", pad=True)
)
assert second.status_code == 201
body = second.json()
# Created, not skipped: a metadata match is a guess, and refusing a legitimate
# second edition costs more than a note does.
assert len(body["created"]) == 1
assert body["skipped"] == []
assert len(body["possible_duplicates"]) == 1
possible = body["possible_duplicates"][0]
assert possible["book_id"] == body["created"][0]["id"]
candidate = possible["candidates"][0]
assert candidate["book_id"] == original["id"]
assert candidate["title"] == original["title"]
assert candidate["authors"] == ["Franz Kafka"]
assert "title-author" in candidate["matched_on"]
async def test_the_review_screen_groups_them(
self, authenticated_client: AsyncClient
) -> None:
first = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("first.epub")
)
second = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("second.epub", pad=True)
)
book_ids = sorted(
[first.json()["created"][0]["id"], second.json()["created"][0]["id"]]
)
response = await authenticated_client.get("/books/duplicate-books?library_id=1")
assert response.status_code == 200
groups = response.json()
assert len(groups) == 1
assert [book["book_id"] for book in groups[0]["books"]] == book_ids
async def test_a_dismissed_group_stays_dismissed(
self, authenticated_client: AsyncClient
) -> None:
first = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("first.epub")
)
second = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("second.epub", pad=True)
)
pair = {
"book_a_id": first.json()["created"][0]["id"],
"book_b_id": second.json()["created"][0]["id"],
}
dismissed = await authenticated_client.post(
"/books/duplicate-books/dismissals", json=pair
)
assert dismissed.status_code == 204
response = await authenticated_client.get("/books/duplicate-books?library_id=1")
assert response.json() == []
restored = await authenticated_client.delete(
"/books/duplicate-books/dismissals"
f"?book_a_id={pair['book_b_id']}&book_b_id={pair['book_a_id']}"
)
assert restored.status_code == 204
response = await authenticated_client.get("/books/duplicate-books?library_id=1")
assert len(response.json()) == 1
async def test_two_books_merge_into_one(
self, authenticated_client: AsyncClient
) -> None:
"""The survivor keeps its id and gains the other's file; the other is gone."""
first = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("first.epub")
)
second = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("second.epub", pad=True)
)
keep = first.json()["created"][0]
fold = second.json()["created"][0]
response = await authenticated_client.post(
"/books/merge?library_id=1",
json={
"survivor_id": keep["id"],
"merged_ids": [fold["id"]],
"metadata": {"title": "Metamorphosis", "edition": 2},
},
)
assert response.status_code == 201
merged = response.json()
assert merged["id"] == keep["id"]
assert merged["edition"] == 2
assert len(merged["files"]) == 2
# The folded record is gone, and the group it formed with it.
assert (await authenticated_client.get(f"/books/{fold['id']}")).status_code == 404
groups = await authenticated_client.get("/books/duplicate-books?library_id=1")
assert groups.json() == []
async def test_merging_an_unknown_book_is_refused(
self, authenticated_client: AsyncClient
) -> None:
created = await authenticated_client.post(
"/books/fromFiles?library_id=1", files=self.upload("first.epub")
)
keep = created.json()["created"][0]["id"]
response = await authenticated_client.post(
"/books/merge?library_id=1",
json={"survivor_id": keep, "merged_ids": [9999]},
)
assert response.status_code == 400
async def test_dismissing_an_unknown_book_is_refused(
self, authenticated_client: AsyncClient
) -> None:
response = await authenticated_client.post(
"/books/duplicate-books/dismissals",
json={"book_a_id": 1, "book_b_id": 9999},
)
assert response.status_code == 400
# NOTE: the multi-book ZIP download is covered at the service level, in
# tests/unit/test_services/test_book_service.py. Driving `/books/download` through
# AsyncTestClient hangs in fixture teardown: it is the only `Stream` endpoint in the
# app, and the test transport never sends the `http.disconnect` that Litestar's
# streaming response waits on, so the app's lifespan shutdown never completes.
# async def test_delete_book_metadata(authenticated_client: AsyncClient) -> None:
@@ -868,3 +1272,94 @@ class TestFileManagement:
)
# Should succeed (idempotent)
assert response2.status_code == 204
@pytest.mark.asyncio
class TestUnnameableFormats:
"""
A file whose format nothing can name must still round-trip.
`mimetypes.guess_type` answers None for `.mobi`, `.azw`, `.fb2` and `.lit`, which
is most of what a library imported from elsewhere carries alongside its EPUBs.
`FileMetadataRead.content_type` used to be a required string, so such a book was
created and then failed serialisation on its way back out a 500 on a book the
reader can otherwise download.
"""
def upload(self, name: str) -> list[tuple[str, tuple]]:
# `application/octet-stream` is what a browser posts for these, and it is not
# an answer — the extension is what names the format.
return [("files", (name, b"BOOKMOBI\x00 payload", "application/octet-stream"))]
async def test_a_mobi_is_named_from_its_extension(
self, authenticated_client: AsyncClient
) -> None:
response = await authenticated_client.post(
"/books?library_id=1", files=self.upload("Dune.mobi"), data={"library_id": 1}
)
assert response.status_code == 201
book = response.json()
assert book["files"][0]["content_type"] == "application/x-mobipocket-ebook"
detail = await authenticated_client.get(f"/books/{book['id']}")
assert detail.status_code == 200
async def test_an_unknown_extension_stores_no_content_type(
self, authenticated_client: AsyncClient
) -> None:
"""Null, not a placeholder — and the book still serialises either way."""
response = await authenticated_client.post(
"/books?library_id=1",
files=self.upload("Notes.xyzzy"),
data={"library_id": 1},
)
assert response.status_code == 201
book = response.json()
assert book["files"][0]["content_type"] is None
detail = await authenticated_client.get(f"/books/{book['id']}")
assert detail.status_code == 200
assert detail.json()["files"][0]["content_type"] is None
async def test_the_file_downloads(
self, authenticated_client: AsyncClient
) -> None:
"""Litestar supplies its own media type when the row carries none."""
created = await authenticated_client.post(
"/books?library_id=1",
files=self.upload("Notes.xyzzy"),
data={"library_id": 1},
)
book = created.json()
response = await authenticated_client.get(
f"/books/download/{book['id']}/{book['files'][0]['id']}"
)
assert response.status_code == 200
assert response.headers["content-type"] == "application/octet-stream"
async def test_the_opds_feed_survives_a_null_content_type(
self, authenticated_client: AsyncClient
) -> None:
"""
The one place the type has to be a string.
`Link.type` is required, so a null fails the whole feed rather than one entry.
OPDS clients speak Basic, not the JWT the rest of the API uses.
"""
await authenticated_client.post(
"/books?library_id=1",
files=self.upload("Notes.xyzzy"),
data={"library_id": 1},
)
feed = await authenticated_client.get(
"/opds/acquisition?feed_id=all&feed_title=All+Books",
auth=("user1@example.com", "password123"),
)
assert feed.status_code == 200
assert 'type="application/octet-stream"' in feed.text
@@ -0,0 +1,400 @@
"""
Tests for the Calibre import endpoints.
The API takes a zipped library and nothing else a desktop Calibre install is usually
not on the server, and importing from a path the server can already see stays a
server-side operation (`litestar calibre-import`).
"""
import asyncio
import tempfile
import zipfile
from pathlib import Path
import pytest
from httpx import AsyncClient
from chitai.services.calibre_import import registry
from tests.calibre_fixtures import CalibreFixture
EPUB = Path("tests/data_files/Metamorphosis - Franz Kafka.epub")
OTHER_EPUB = Path("tests/data_files/The Art of War - Sun Tzu.epub")
@pytest.fixture(autouse=True)
def _clear_registry():
"""The registry is a module-level singleton, so it leaks between tests."""
registry._jobs.clear()
yield
registry._jobs.clear()
@pytest.fixture(name="source")
def fx_source(tmp_path: Path) -> Path:
fixture = CalibreFixture(tmp_path / "calibre")
fixture.add_book(
1,
"The Metamorphosis",
authors=["Franz Kafka"],
tags=["Fiction"],
cover=True,
formats={"EPUB": EPUB},
)
fixture.add_book(2, "The Art of War", authors=["Sun Tzu"], formats={"EPUB": OTHER_EPUB})
fixture.add_book(3, "Metadata Only", authors=["Nobody"])
return fixture.commit()
def zip_of(root: Path, into: Path, prefix: str = "") -> Path:
"""Zip a directory the way a file manager would."""
into.mkdir(parents=True, exist_ok=True)
archive = into / "library.zip"
with zipfile.ZipFile(archive, "w") as writing:
for path in sorted(root.rglob("*")):
if path.is_file():
writing.write(path, f"{prefix}{path.relative_to(root)}")
return archive
async def upload(
client: AsyncClient,
archive: Path,
library_id: int = 1,
allow_duplicates: bool = False,
) -> tuple[int, dict]:
response = await client.post(
f"/libraries/{library_id}/imports/calibre/upload",
files=[("archive", (archive.name, archive.read_bytes(), "application/zip"))],
data={"allow_duplicates": str(allow_duplicates).lower()},
)
return response.status_code, response.json()
async def wait_for(client: AsyncClient, job_id: str) -> dict:
"""Poll until the job is no longer running, the way the screen does."""
for _ in range(200):
response = await client.get(f"/libraries/imports/{job_id}")
assert response.status_code == 200
job = response.json()
if job["state"] != "running":
return job
await asyncio.sleep(0.05)
raise AssertionError("the import never finished")
async def test_an_uploaded_library_imports(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
status, job = await upload(
authenticated_client, zip_of(source, tmp_path / "out", prefix="Calibre Library/")
)
assert status == 202
assert job["state"] == "running"
assert job["library_id"] == 1
# The archive's name, not the temp directory it was unpacked into.
assert job["source"] == "library.zip"
finished = await wait_for(authenticated_client, job["id"])
assert finished["state"] == "finished"
assert finished["total"] == 3
assert finished["created"] == 2
assert finished["skipped"] == 1
assert finished["failed"] == 0
assert finished["error"] is None
assert finished["current_title"] is None
listed = await authenticated_client.get("/books?library_id=1")
titles = [book["title"] for book in listed.json()["items"]]
assert "The Metamorphosis" in titles
assert "The Art of War" in titles
async def test_a_library_zipped_without_a_wrapping_folder(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
"""Zipping the contents is as common as zipping the folder."""
status, job = await upload(authenticated_client, zip_of(source, tmp_path / "out"))
assert status == 202
finished = await wait_for(authenticated_client, job["id"])
assert finished["created"] == 2
async def test_the_unpacked_copy_is_cleaned_up(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
"""
An unpacked archive is a second copy of the whole library.
The books worth keeping have been copied into the library by the time the job ends,
so nothing is lost with it and nothing will come back for it.
"""
status, job = await upload(authenticated_client, zip_of(source, tmp_path / "out"))
assert status == 202
workspace = registry.get(job["id"]).workspace
assert workspace is not None
await wait_for(authenticated_client, job["id"])
assert not workspace.exists()
async def test_uploading_the_same_library_twice_imports_nothing_new(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
"""Re-running is safe, which is what makes an interrupted import resumable."""
archive = zip_of(source, tmp_path / "out")
for _ in range(2):
_, job = await upload(authenticated_client, archive)
finished = await wait_for(authenticated_client, job["id"])
assert finished["created"] == 0
assert finished["skipped"] == 3
async def test_allow_duplicates_stores_the_files_again(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
"""
The one option the screen offers, and it has to reach the import.
Without it the second pass skips everything, which is the previous test.
"""
archive = zip_of(source, tmp_path / "out")
_, first = await upload(authenticated_client, archive)
await wait_for(authenticated_client, first["id"])
_, second = await upload(authenticated_client, archive, allow_duplicates=True)
finished = await wait_for(authenticated_client, second["id"])
assert finished["created"] == 2
assert finished["skipped"] == 1 # still the book with no files
async def test_two_imports_into_one_library_conflict(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
archive = zip_of(source, tmp_path / "out")
first_status, first = await upload(authenticated_client, archive)
assert first_status == 202
second_status, second = await upload(authenticated_client, archive)
assert second_status == 409
assert second["extra"]["job_id"] == first["id"]
await wait_for(authenticated_client, first["id"])
async def test_a_finished_import_does_not_block_the_next_one(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
archive = zip_of(source, tmp_path / "out")
_, first = await upload(authenticated_client, archive)
await wait_for(authenticated_client, first["id"])
status, second = await upload(authenticated_client, archive)
assert status == 202
await wait_for(authenticated_client, second["id"])
async def test_cancelling_stops_after_the_current_book(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
"""
Cancelling is not aborting: a book abandoned mid-copy would leave files with no row.
Whether this cancels before any book, after one, or after the lot is a race the
catalogue is three books long. What must hold either way is that the state is
terminal and every book it did import is complete.
"""
_, job = await upload(authenticated_client, zip_of(source, tmp_path / "out"))
cancelled = await authenticated_client.delete(f"/libraries/imports/{job['id']}")
assert cancelled.status_code == 200
final = await wait_for(authenticated_client, job["id"])
assert final["state"] in {"cancelled", "finished"}
listed = await authenticated_client.get("/books?library_id=1")
for book in listed.json()["items"]:
assert book["files"]
async def test_failures_are_reported_on_the_job(
authenticated_client: AsyncClient, tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""One book failing is recorded and does not stop the run."""
fixture = CalibreFixture(tmp_path / "calibre")
fixture.add_book(1, "Fine", authors=["A"], formats={"EPUB": EPUB})
fixture.add_book(2, "Doomed", authors=["B"], formats={"EPUB": OTHER_EPUB})
root = fixture.commit()
from chitai.services.book import BookService
original = BookService.create
async def fail_on_the_second(self, data, *args, **kwargs):
if isinstance(data, dict) and data.get("title") == "Doomed":
raise RuntimeError("no room on the shelf")
return await original(self, data, *args, **kwargs)
monkeypatch.setattr(BookService, "create", fail_on_the_second)
_, job = await upload(authenticated_client, zip_of(root, tmp_path / "out"))
finished = await wait_for(authenticated_client, job["id"])
assert finished["state"] == "finished"
assert finished["created"] == 1
assert finished["failed"] == 1
assert finished["failures"][0]["calibre_id"] == 2
assert "no room on the shelf" in finished["failures"][0]["reason"]
async def test_a_second_copy_is_counted_as_a_possible_duplicate(
authenticated_client: AsyncClient, tmp_path: Path
) -> None:
"""A count, not the records — the duplicates screen is what shows them."""
padded = tmp_path / "padded.epub"
padded.write_bytes(EPUB.read_bytes() + b"\0" * 64)
fixture = CalibreFixture(tmp_path / "calibre")
fixture.add_book(1, "The Metamorphosis", authors=["Franz Kafka"], formats={"EPUB": EPUB})
fixture.add_book(
2, "The Metamorphosis", authors=["Franz Kafka"], formats={"EPUB": padded}
)
root = fixture.commit()
_, job = await upload(authenticated_client, zip_of(root, tmp_path / "out"))
finished = await wait_for(authenticated_client, job["id"])
assert finished["created"] == 2
assert finished["possible_duplicates"] == 1
async def test_polling_an_unknown_job(authenticated_client: AsyncClient) -> None:
response = await authenticated_client.get("/libraries/imports/not-a-job")
assert response.status_code == 404
async def test_cancelling_an_unknown_job(authenticated_client: AsyncClient) -> None:
response = await authenticated_client.delete("/libraries/imports/not-a-job")
assert response.status_code == 404
async def test_importing_into_a_library_that_does_not_exist(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
status, _ = await upload(
authenticated_client, zip_of(source, tmp_path / "out"), library_id=999
)
assert status == 404
async def test_importing_into_a_read_only_library(
authenticated_client: AsyncClient, source: Path, tmp_path: Path
) -> None:
"""A read-only library points at a tree Chitai does not own."""
root = tmp_path / "read-only"
root.mkdir()
created = await authenticated_client.post(
"/libraries",
json={"name": "Read Only", "root_path": str(root), "read_only": True},
)
assert created.status_code == 201
status, body = await upload(
authenticated_client,
zip_of(source, tmp_path / "out"),
library_id=created.json()["id"],
)
assert status == 400
assert "read-only" in body["detail"]
async def test_an_import_needs_authentication(
client: AsyncClient, source: Path, tmp_path: Path
) -> None:
status, _ = await upload(client, zip_of(source, tmp_path / "out"))
assert status == 401
class TestRefusedArchives:
"""Everything wrong with an archive is answered now, not as a job that fails later."""
async def test_a_hostile_archive(
self, authenticated_client: AsyncClient, tmp_path: Path
) -> None:
"""Zip slip."""
archive = tmp_path / "hostile.zip"
with zipfile.ZipFile(archive, "w") as writing:
writing.writestr("metadata.db", "not really")
writing.writestr("../../escaped.txt", "gotcha")
status, body = await upload(authenticated_client, archive)
assert status == 400
assert "outside itself" in body["detail"]
assert registry._jobs == {}
async def test_something_that_is_not_a_zip(
self, authenticated_client: AsyncClient, tmp_path: Path
) -> None:
archive = tmp_path / "notes.txt"
archive.write_bytes(b"just some text")
status, body = await upload(authenticated_client, archive)
assert status == 400
assert "not a zip" in body["detail"]
async def test_an_archive_with_no_catalogue(
self, authenticated_client: AsyncClient, tmp_path: Path
) -> None:
archive = tmp_path / "books.zip"
with zipfile.ZipFile(archive, "w") as writing:
writing.writestr("Some Book.epub", "content")
status, body = await upload(authenticated_client, archive)
assert status == 400
assert "no metadata.db" in body["detail"]
async def test_a_refusal_leaves_no_temp_files(
self, authenticated_client: AsyncClient, tmp_path: Path
) -> None:
"""Every refusal path removes the workspace it had already made."""
before = set(Path(tempfile.gettempdir()).glob("tmp*"))
archive = tmp_path / "books.zip"
with zipfile.ZipFile(archive, "w") as writing:
writing.writestr("Some Book.epub", "content")
await upload(authenticated_client, archive)
assert set(Path(tempfile.gettempdir()).glob("tmp*")) == before
@@ -0,0 +1,98 @@
"""Tests for KOReader-compatible file hash generation."""
import pytest
from httpx import AsyncClient
from pathlib import Path
# Known KOReader hashes for test files
TEST_FILES = {
"Moby Dick; Or, The Whale - Herman Melville.epub": {
"path": Path("tests/data_files/Moby Dick; Or, The Whale - Herman Melville.epub"),
"hash": "ceeef909ec65653ba77e1380dff998fb",
"content_type": "application/epub+zip",
},
"Calculus Made Easy - Silvanus Thompson.pdf": {
"path": Path("tests/data_files/Calculus Made Easy - Silvanus Thompson.pdf"),
"hash": "ace67d512efd1efdea20f3c2436b6075",
"content_type": "application/pdf",
},
}
@pytest.mark.parametrize(
("book_name",),
[(name,) for name in TEST_FILES.keys()],
)
async def test_upload_book_generates_correct_hash(
authenticated_client: AsyncClient,
book_name: str,
) -> None:
"""Test that uploading a book generates the correct KOReader-compatible hash."""
book_info = TEST_FILES[book_name]
file_content = book_info["path"].read_bytes()
files = [("files", (book_name, file_content, book_info["content_type"]))]
data = {"library_id": "1"}
response = await authenticated_client.post(
"/books?library_id=1",
files=files,
data=data,
)
assert response.status_code == 201
book_data = response.json()
assert len(book_data["files"]) == 1
file_metadata = book_data["files"][0]
assert "hash" in file_metadata
assert file_metadata["hash"] == book_info["hash"]
async def test_add_file_to_book_generates_correct_hash(
authenticated_client: AsyncClient,
) -> None:
"""Test that adding a file to an existing book generates the correct hash."""
# Create a book with the first file
first_book = TEST_FILES["Moby Dick; Or, The Whale - Herman Melville.epub"]
first_content = first_book["path"].read_bytes()
files = [("files", (first_book["path"].name, first_content, first_book["content_type"]))]
data = {"library_id": "1"}
create_response = await authenticated_client.post(
"/books?library_id=1",
files=files,
data=data,
)
assert create_response.status_code == 201
book_id = create_response.json()["id"]
# Add the second file to the book
second_book = TEST_FILES["Calculus Made Easy - Silvanus Thompson.pdf"]
second_content = second_book["path"].read_bytes()
add_files = [("data", (second_book["path"].name, second_content, second_book["content_type"]))]
add_response = await authenticated_client.post(
f"/books/{book_id}/files",
files=add_files,
)
assert add_response.status_code == 201
updated_book = add_response.json()
# Verify both files have correct hashes
assert len(updated_book["files"]) == 2
for file_metadata in updated_book["files"]:
assert "hash" in file_metadata
epub_file = next(f for f in updated_book["files"] if f["path"].endswith(".epub"))
pdf_file = next(f for f in updated_book["files"] if f["path"].endswith(".pdf"))
assert epub_file["hash"] == first_book["hash"]
assert pdf_file["hash"] == second_book["hash"]
+392
View File
@@ -0,0 +1,392 @@
import zipfile
from datetime import date
from pathlib import Path
from types import SimpleNamespace
import pytest
from chitai.services.calibre import (
CalibreLibrary,
CalibreLibraryError,
extract_calibre_archive,
format_series_index,
parse_date,
strip_html,
unescape_author,
)
from tests.calibre_fixtures import UNDEFINED_DATE, CalibreFixture
EPUB = Path("tests/data_files/Metamorphosis - Franz Kafka.epub")
PDF = Path("tests/data_files/Calculus Made Easy - Silvanus Thompson.pdf")
@pytest.fixture(name="library_root")
def fx_library_root(tmp_path: Path) -> Path:
"""A small Calibre library covering the rows that are easy to read wrongly."""
fixture = CalibreFixture(tmp_path / "Calibre Library")
fixture.add_pages_table()
fixture.add_book(
1,
"The Metamorphosis",
authors=["Franz Kafka"],
pubdate="1915-10-15 00:00:00+00:00",
tags=["Fiction", "Absurdist"],
publisher="Kurt Wolff Verlag",
languages=["deu", "eng"],
comment="<p>A travelling salesman.</p><p>He wakes up <i>changed</i>.</p>",
identifiers={"isbn": "978-0-486-29030-0", "amazon": "B01N5IB20Q"},
uuid="11111111-2222-3333-4444-555555555555",
pages=201,
cover=True,
formats={"EPUB": EPUB},
)
# Volume seven of a series, and no publication date — the two values most likely to
# be carried through verbatim when they should not be.
fixture.add_book(
2,
"Persepolis Rising",
# Calibre escapes the comma and nothing else, so the space after it is stored
# as-is: `Corey, Jr.` is written `Corey| Jr.`.
authors=["Corey| Jr., James S. A."],
series="The Expanse",
series_index=7.0,
formats={"EPUB": EPUB},
)
# A novella between two novels: a fractional position is real and must survive.
fixture.add_book(
3, "Strange Dogs", series="The Expanse", series_index=6.5, formats={"PDF": PDF}
)
# Every row Calibre will happily hold and Chitai cannot use: no files at all.
fixture.add_book(4, "Metadata Only")
# A catalogue row whose file is not on disk.
fixture.add_book(5, "Lost Book")
fixture.add_missing_format(5, "EPUB", "Lost Book - Unknown")
return fixture.commit()
async def test_reads_a_book_whole(library_root: Path) -> None:
async with CalibreLibrary(library_root) as library:
assert await library.count() == 5
books = await library.books()
book = books[0]
assert book.calibre_id == 1
assert book.title == "The Metamorphosis"
assert book.authors == ["Franz Kafka"]
assert book.published_date == date(1915, 10, 15)
assert book.tags == ["Absurdist", "Fiction"]
assert book.publisher == "Kurt Wolff Verlag"
assert book.pages == 201
assert book.uuid == "11111111-2222-3333-4444-555555555555"
# One language, and the one Calibre put first.
assert book.language == "deu"
# Reported as Calibre wrote them: folding `amazon` onto `asin` is the importer's
# job, not the reader's.
assert book.identifiers == {"isbn": "978-0-486-29030-0", "amazon": "B01N5IB20Q"}
assert book.cover is not None
assert book.cover.is_file()
assert len(book.files) == 1
assert book.files[0].format == "EPUB"
assert book.files[0].path.is_file()
# The stem is Calibre's, truncated and sanitised — never the title.
assert book.files[0].path.name != f"{book.title}.epub"
async def test_the_undefined_date_is_not_a_date(library_root: Path) -> None:
"""`0101-01-01` parses fine, which is exactly the problem."""
async with CalibreLibrary(library_root) as library:
books = {book.calibre_id: book for book in await library.books()}
assert books[2].published_date is None
async def test_series_position_is_a_plain_string(library_root: Path) -> None:
async with CalibreLibrary(library_root) as library:
books = {book.calibre_id: book for book in await library.books()}
assert books[2].series == "The Expanse"
assert books[2].series_position == "7"
assert books[3].series_position == "6.5"
# `series_index` defaults to 1.0 for every book, so a position without a series
# would invent a volume one out of nothing.
assert books[1].series is None
assert books[1].series_position is None
async def test_author_commas_are_unescaped(library_root: Path) -> None:
async with CalibreLibrary(library_root) as library:
books = {book.calibre_id: book for book in await library.books()}
assert books[2].authors == ["Corey, Jr., James S. A."]
async def test_comments_come_back_as_text(library_root: Path) -> None:
async with CalibreLibrary(library_root) as library:
books = {book.calibre_id: book for book in await library.books()}
assert books[1].description == "A travelling salesman.\nHe wakes up changed."
assert books[2].description is None
async def test_files_are_reported_whether_or_not_they_exist(library_root: Path) -> None:
"""
The reader says what the catalogue says. Whether the bytes are there is a question
for whoever is about to copy them, which stats them anyway.
"""
async with CalibreLibrary(library_root) as library:
books = {book.calibre_id: book for book in await library.books()}
assert books[4].files == []
assert len(books[5].files) == 1
assert not books[5].files[0].path.exists()
async def test_a_library_without_the_pages_table_still_reads(tmp_path: Path) -> None:
"""`books_pages_link` is recent; an older library simply does not have it."""
fixture = CalibreFixture(tmp_path / "Old Library")
fixture.add_book(1, "Old Book", formats={"EPUB": EPUB})
root = fixture.commit()
async with CalibreLibrary(root) as library:
books = await library.books()
assert books[0].pages is None
async def test_the_original_is_never_opened(library_root: Path) -> None:
"""
The catalogue is copied before it is read, and the copy goes away afterwards.
Calibre may be running and writing; this is what keeps a live library out of it.
"""
before = (library_root / "metadata.db").read_bytes()
library = CalibreLibrary(library_root)
await library.open()
workspace = library._workspace
assert workspace is not None and (workspace / "metadata.db").is_file()
await library.close()
assert not workspace.exists()
assert (library_root / "metadata.db").read_bytes() == before
async def test_closing_twice_is_harmless(library_root: Path) -> None:
library = CalibreLibrary(library_root)
await library.open()
await library.close()
await library.close()
async def test_a_directory_that_is_not_a_calibre_library(tmp_path: Path) -> None:
with pytest.raises(CalibreLibraryError, match="not a Calibre library"):
await CalibreLibrary(tmp_path).open()
@pytest.mark.parametrize(
("stored", "expected"),
[
("2017-12-04 04:00:00+00:00", date(2017, 12, 4)),
("2001-07-02 00:00:00+00:00", date(2001, 7, 2)),
("1999-01-31", date(1999, 1, 31)),
# Calibre's sentinel, and anything else implausibly early.
(UNDEFINED_DATE, None),
("0101-01-01", None),
(None, None),
("", None),
("not a date", None),
],
)
def test_parse_date(stored: str | None, expected: date | None) -> None:
assert parse_date(stored) == expected
@pytest.mark.parametrize(
("index", "expected"),
[
(7.0, "7"),
(1.0, "1"),
(6.5, "6.5"),
(0.0, "0"),
(12.25, "12.25"),
(None, None),
],
)
def test_format_series_index(index: float | None, expected: str | None) -> None:
assert format_series_index(index) == expected
@pytest.mark.parametrize(
("stored", "expected"),
[
("Doyle| Sir Arthur Conan", "Doyle, Sir Arthur Conan"),
("Franz Kafka", "Franz Kafka"),
(" Herman Melville ", "Herman Melville"),
],
)
def test_unescape_author(stored: str, expected: str) -> None:
assert unescape_author(stored) == expected
@pytest.mark.parametrize(
("html", "expected"),
[
("<p>One.</p><p>Two.</p>", "One.\nTwo."),
("Plain text", "Plain text"),
("<div>A<br>B</div>", "A\nB"),
("<p>Caf&eacute; &amp; bar</p>", "Café & bar"),
("<ul><li>One</li><li>Two</li></ul>", "One\nTwo"),
# Markup carrying no text at all is nothing, not an empty description.
("<p></p>", None),
("", None),
(None, None),
],
)
def test_strip_html(html: str | None, expected: str | None) -> None:
assert strip_html(html) == expected
class TestArchives:
"""A Calibre library that arrives zipped rather than as a path."""
def zipped(self, root: Path, into: Path, prefix: str = "") -> Path:
"""Zip a directory the way a file manager would."""
archive = into / "library.zip"
with zipfile.ZipFile(archive, "w") as writing:
for path in sorted(root.rglob("*")):
if path.is_file():
writing.write(path, f"{prefix}{path.relative_to(root)}")
return archive
async def test_a_library_zipped_at_its_root(
self, library_root: Path, tmp_path: Path
) -> None:
archive = self.zipped(library_root, tmp_path)
destination = tmp_path / "unpacked"
destination.mkdir()
catalogue = await extract_calibre_archive(archive, destination)
assert catalogue == destination
async with CalibreLibrary(catalogue) as library:
assert await library.count() == 5
async def test_a_library_zipped_inside_a_folder(
self, library_root: Path, tmp_path: Path
) -> None:
"""Zipping the folder itself is at least as common as zipping its contents."""
archive = self.zipped(library_root, tmp_path, prefix="Calibre Library/")
destination = tmp_path / "unpacked"
destination.mkdir()
catalogue = await extract_calibre_archive(archive, destination)
assert catalogue == destination / "Calibre Library"
async with CalibreLibrary(catalogue) as library:
assert await library.count() == 5
async def test_an_entry_pointing_outside_the_archive_is_refused(
self, tmp_path: Path
) -> None:
"""
Zip slip. `ZipFile.extract` sanitises names itself, but relying on that silently
is how the next person to change the extraction call reintroduces it.
"""
archive = tmp_path / "hostile.zip"
with zipfile.ZipFile(archive, "w") as writing:
writing.writestr("metadata.db", "not really")
writing.writestr("../../escaped.txt", "gotcha")
destination = tmp_path / "unpacked"
destination.mkdir()
with pytest.raises(CalibreLibraryError, match="outside itself"):
await extract_calibre_archive(archive, destination)
assert not (tmp_path.parent / "escaped.txt").exists()
async def test_something_that_is_not_a_zip(self, tmp_path: Path) -> None:
archive = tmp_path / "not.zip"
archive.write_bytes(b"PK-ish, but no")
destination = tmp_path / "unpacked"
destination.mkdir()
with pytest.raises(CalibreLibraryError, match="not a zip file"):
await extract_calibre_archive(archive, destination)
async def test_an_archive_with_no_catalogue(self, tmp_path: Path) -> None:
archive = tmp_path / "books.zip"
with zipfile.ZipFile(archive, "w") as writing:
writing.writestr("Some Book.epub", "content")
destination = tmp_path / "unpacked"
destination.mkdir()
with pytest.raises(CalibreLibraryError, match="no metadata.db"):
await extract_calibre_archive(archive, destination)
# Refused before anything was written.
assert list(destination.iterdir()) == []
async def test_a_catalogue_buried_too_deep(self, tmp_path: Path) -> None:
"""Somebody's whole backup tree is not a library, however much it contains one."""
archive = tmp_path / "backup.zip"
with zipfile.ZipFile(archive, "w") as writing:
writing.writestr("backups/2026/january/library/metadata.db", "not really")
destination = tmp_path / "unpacked"
destination.mkdir()
with pytest.raises(CalibreLibraryError, match="within 3 levels"):
await extract_calibre_archive(archive, destination)
async def test_an_archive_too_big_for_the_disk(
self, tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""
Checked before writing, not discovered part-way through.
A full disk takes the whole application down, and the size is in the archive
already.
"""
archive = tmp_path / "huge.zip"
with zipfile.ZipFile(archive, "w") as writing:
writing.writestr("metadata.db", "not really")
destination = tmp_path / "unpacked"
destination.mkdir()
monkeypatch.setattr(
"chitai.services.calibre.shutil.disk_usage",
lambda _path: SimpleNamespace(total=1024, used=1024, free=0),
)
with pytest.raises(CalibreLibraryError, match="only"):
await extract_calibre_archive(archive, destination)
assert list(destination.iterdir()) == []
+64
View File
@@ -0,0 +1,64 @@
from pathlib import Path
import pytest
from chitai.services.utils import guess_content_type
@pytest.mark.parametrize(
("filename", "expected"),
[
# What `mimetypes` already knows, kept here so a host with a thin
# /etc/mime.types cannot change the answer without a test noticing.
("Frankenstein.epub", "application/epub+zip"),
("Calculus.pdf", "application/pdf"),
("Persepolis.azw3", "application/vnd.amazon.mobi8-ebook"),
("Watchmen.cbz", "application/vnd.comicbook+zip"),
# What it does not, and where a Calibre library's older formats live.
("Dune.mobi", "application/x-mobipocket-ebook"),
("Dune.prc", "application/x-mobipocket-ebook"),
("Dune.azw", "application/vnd.amazon.ebook"),
("Voyna i Mir.fb2", "application/x-fictionbook+xml"),
("Voyna i Mir.fbz", "application/x-zip-compressed-fb2"),
("Reader.lit", "application/x-ms-reader"),
("Reader.lrf", "application/x-sony-bbeb"),
("Watchmen.cb7", "application/x-cb7"),
# Case is not part of the answer, and Calibre writes formats uppercase.
("Dune.MOBI", "application/x-mobipocket-ebook"),
# Nothing can name these, and None is the answer rather than a placeholder.
("Notes.xyzzy", None),
("README", None),
],
)
def test_guess_content_type(filename: str, expected: str | None) -> None:
assert guess_content_type(Path(filename)) == expected
# A str and a Path must agree, and an upload's `filename` carries its relative
# path, so a name with directories in front of it has to resolve the same way.
assert guess_content_type(filename) == expected
assert guess_content_type(f"Some Author/Some Book/{filename}") == expected
def test_fallback_is_used_only_when_the_extension_says_nothing() -> None:
"""A client's claim fills a gap; it never overrides the name."""
assert (
guess_content_type(Path("Dune.mobi"), fallback="application/pdf")
== "application/x-mobipocket-ebook"
)
assert (
guess_content_type(Path("Notes.xyzzy"), fallback="application/epub+zip")
== "application/epub+zip"
)
def test_an_unspecified_fallback_is_not_an_answer() -> None:
"""
`application/octet-stream` from a client is it saying it does not know.
Browsers post exactly that for every extension they do not recognise, which is most
ebook formats. Storing it would be indistinguishable from having determined a
format, so it is discarded and the column keeps its null.
"""
assert (
guess_content_type(Path("Notes.xyzzy"), fallback="application/octet-stream")
is None
)
@@ -0,0 +1,69 @@
"""Tests for BookPathGenerator."""
from pathlib import Path
from chitai.services.filesystem_library import BookPathGenerator, sanitize_path_component
ROOT = Path("/library")
def path_for(**book) -> Path:
return BookPathGenerator(ROOT).generate_path(book)
def test_author_and_title() -> None:
assert path_for(title="Dune", authors=["Frank Herbert"]) == (
ROOT / "Frank Herbert" / "Dune"
)
def test_a_book_with_no_authors() -> None:
assert path_for(title="Beowulf", authors=[]) == ROOT / "Unknown" / "Beowulf"
def test_a_series_adds_a_level_and_pads_the_position() -> None:
assert path_for(
title="Persepolis Rising",
authors=["James S. A. Corey"],
series="The Expanse",
series_position="7",
) == ROOT / "James S. A. Corey" / "The Expanse" / "07 - Persepolis Rising"
def test_a_slash_in_a_title_does_not_add_a_directory() -> None:
"""
The separators in the path come from the template, never from the metadata.
A title with a slash in it "AC/DC", "Him/Her" would otherwise put the book one
level below where `book.path` says it is, which is what deletes, moves and file
lookups all act on. Calibre keeps the real title in its database and strips this
from its own directory names, so an import is where they surface.
"""
generated = path_for(title="Back in Black: AC/DC", authors=["Murray Engleheart"])
assert generated == ROOT / "Murray Engleheart" / "Back in Black: AC_DC"
assert generated.relative_to(ROOT).parts == ("Murray Engleheart", "Back in Black: AC_DC")
def test_a_slash_in_an_author_or_series_is_handled_too() -> None:
assert path_for(title="Split", authors=["A/B Collective"]) == (
ROOT / "A_B Collective" / "Split"
)
assert path_for(
title="Volume One", authors=["Someone"], series="Either/Or", series_position="1"
) == ROOT / "Someone" / "Either_Or" / "01 - Volume One"
def test_control_characters_are_removed() -> None:
assert path_for(title="Line\nBreak", authors=["Someone"]) == (
ROOT / "Someone" / "Line_Break"
)
def test_sanitize_path_component() -> None:
assert sanitize_path_component("AC/DC") == "AC_DC"
assert sanitize_path_component("back\\slash") == "back_slash"
assert sanitize_path_component(" padded ") == "padded"
# Colons and other punctuation are legal in a path and are left alone.
assert sanitize_path_component("Title: Subtitle") == "Title: Subtitle"
+233
View File
@@ -0,0 +1,233 @@
"""Tests for the normalization behind book-level duplicate detection."""
import pytest
from chitai.services.matching import (
format_author_name,
normalize_author,
normalize_identifier,
normalize_title,
)
from chitai.services.metadata_extractor import parse_identifier
from chitai.services.utils import isbn10_to_isbn13
class TestNormalizeTitle:
"""Two copies of one book rarely agree on how the title is written."""
@pytest.mark.parametrize(
("title", "expected"),
[
("The Metamorphosis", "metamorphosis"),
("Metamorphosis", "metamorphosis"),
("METAMORPHOSIS", "metamorphosis"),
("A Tale of Two Cities", "tale of two cities"),
("An Enquiry", "enquiry"),
# Accents, punctuation and ampersands are spelling, not identity.
("Les Misérables", "les miserables"),
("Moby Dick; Or, The Whale", "moby dick or the whale"),
("Sense & Sensibility", "sense and sensibility"),
# Bracketed asides and trailing edition noise say nothing about the book.
("Frankenstein (Illustrated)", "frankenstein"),
("Frankenstein [Kindle Edition]", "frankenstein"),
("Frankenstein, 2nd Edition", "frankenstein"),
# The compact forms a cover actually carries.
("Building Microservices, 2E", "building microservices"),
("Building Microservices 2e", "building microservices"),
("Frankenstein 3 Ed", "frankenstein"),
("Dungeons & Dragons 5e", "dungeons and dragons"),
("Frankenstein Revised Edition", "frankenstein"),
("Dune Deluxe Edition Illustrated", "dune"),
("", ""),
(None, ""),
],
)
def test_titles_that_should_agree(self, title: str | None, expected: str) -> None:
assert normalize_title(title) == expected
def test_a_qualifier_that_is_the_title_survives(self) -> None:
"""A trailing qualifier is noise; the same word at the front is the book."""
assert normalize_title("The Illustrated Man") == "illustrated man"
def test_normalization_never_empties_a_title(self) -> None:
"""An article-only title is not improved by having no article left."""
assert normalize_title("The") == "the"
@pytest.mark.parametrize(
"title", ["Catch 22", "Fahrenheit 451", "Blade Runner 2049", "1984", "Apollo 13"]
)
def test_a_number_is_not_an_edition(self, title: str) -> None:
"""Edition stripping keys on the `e`; a bare number is part of the title."""
assert normalize_title(title) == title.casefold()
class TestNormalizeAuthor:
"""One person, written down several ways."""
@pytest.mark.parametrize(
("name", "expected"),
[
("Franz Kafka", "franz kafka"),
("Kafka, Franz", "franz kafka"),
("KAFKA, FRANZ", "franz kafka"),
("Émile Zola", "emile zola"),
("Doyle, Arthur Conan", "arthur conan doyle"),
# Runs of initials are joined, so spacing them out changes nothing.
("J.R.R. Tolkien", "jrr tolkien"),
("J. R. R. Tolkien", "jrr tolkien"),
("JRR Tolkien", "jrr tolkien"),
("", ""),
(None, ""),
],
)
def test_names_that_should_agree(self, name: str | None, expected: str) -> None:
assert normalize_author(name) == expected
def test_two_people_are_not_reduced_together(self) -> None:
"""Surname plus initial would collide unrelated writers; it is not used."""
assert normalize_author("Charles Dickens") != normalize_author("Colin Dexter")
def test_a_list_is_left_alone(self) -> None:
"""More than one comma is a list or a suffix, and guessing does more harm."""
assert normalize_author("Smith, John, Jr.") == "smith john jr"
class TestFormatAuthorName:
"""What gets stored and shown, as opposed to what gets compared."""
@pytest.mark.parametrize(
("written", "expected"),
[
# A leftover separator from a `DC:creator` list.
("Newman, Sam;", "Sam Newman"),
("Sam Newman ", "Sam Newman"),
(" Dan Vanderkam ", "Dan Vanderkam"),
# `Surname, Given` is how EPUBs file a name, not how anyone reads it.
("Kleppmann, Martin", "Martin Kleppmann"),
("Huxley, Aldous", "Aldous Huxley"),
("Liu, Cixin", "Cixin Liu"),
# An extension carried in from the filename the name was read out of.
("Sam Newman.epub", "Sam Newman"),
("Franz Kafka.mobi", "Franz Kafka"),
("Brian W. Kernighan.epub", "Brian W. Kernighan"),
("", ""),
(None, ""),
],
)
def test_names_are_tidied(self, written: str | None, expected: str) -> None:
assert format_author_name(written) == expected
@pytest.mark.parametrize(
"written",
[
# Two people in one string. Flipping it would invent a third.
"Dave Thomas, Andy Hunt",
"Mark Richards, Neal Ford",
# A compound surname is not recognised, and is left alone rather than
# rearranged wrongly.
"García Márquez, Gabriel",
],
)
def test_an_unrecognised_form_is_left_alone(self, written: str) -> None:
assert format_author_name(written) == written
def test_case_and_accents_belong_to_the_author(self) -> None:
"""Tidying removes what an extractor added; it does not correct spelling."""
assert format_author_name("Michał Płachta.epub") == "Michał Płachta"
assert format_author_name("Steve McConnell") == "Steve McConnell"
@pytest.mark.parametrize(
"written", ["Newman, Sam;", "Sam Newman.epub", "Kleppmann, Martin"]
)
def test_tidying_is_idempotent(self, written: str) -> None:
"""`unique_filter` tidies a name that may already be tidy; it must not drift."""
once = format_author_name(written)
assert format_author_name(once) == once
class TestNormalizeIdentifier:
"""Identifiers only help if the same edition produces the same key."""
def test_isbn_10_and_isbn_13_are_one_key(self) -> None:
assert normalize_identifier("isbn-10", "0486282112") == "isbn:9780486282114"
assert normalize_identifier("isbn-13", "9780486282114") == "isbn:9780486282114"
@pytest.mark.parametrize(
"written", ["978-0-486-28211-4", "978 0 486 28211 4", "9780486282114"]
)
def test_formatting_is_not_part_of_an_isbn(self, written: str) -> None:
assert normalize_identifier("isbn", written) == "isbn:9780486282114"
def test_an_isbn_that_fails_its_checksum_is_no_evidence(self) -> None:
assert normalize_identifier("isbn-13", "9780486282115") is None
def test_uuids_are_refused(self) -> None:
"""Generated per build, so they only re-find what the hash check catches."""
assert normalize_identifier("uuid", "3f2b1c4e-1111-2222-3333-444455556666") is None
assert normalize_identifier("urn:uuid", "3f2b1c4e-1111-2222-3333-444455556666") is None
def test_other_schemes_keep_their_own_key(self) -> None:
assert normalize_identifier("asin", "B000FC0PDA") == "asin:b000fc0pda"
assert normalize_identifier("ASIN", "b000fc0pda") == "asin:b000fc0pda"
def test_something_too_short_is_not_evidence(self) -> None:
"""A Calibre id of "42" would otherwise pair two unrelated books."""
assert normalize_identifier("calibre", "42") is None
@pytest.mark.parametrize(("name", "value"), [("", "1234567"), ("asin", "")])
def test_half_an_identifier_is_no_identifier(self, name: str, value: str) -> None:
assert normalize_identifier(name, value) is None
class TestIsbnConversion:
def test_isbn_10_converts_to_its_isbn_13(self) -> None:
assert isbn10_to_isbn13("0486282112") == "9780486282114"
def test_a_trailing_x_is_a_digit(self) -> None:
assert isbn10_to_isbn13("043942089X") == "9780439420891"
@pytest.mark.parametrize("isbn", ["0486282113", "9780486282114", "nonsense"])
def test_anything_that_is_not_an_isbn_10_converts_to_nothing(self, isbn: str) -> None:
assert isbn10_to_isbn13(isbn) is None
class TestParseIdentifier:
"""What an EPUB writes, and what is worth storing for it."""
@pytest.mark.parametrize(
"written",
[
"9780486282114",
"978-0-486-28211-4",
"urn:isbn:9780486282114",
"urn:isbn:978-0-486-28211-4",
"ISBN:978-0-486-28211-4",
],
)
def test_isbns_survive_however_they_are_written(self, written: str) -> None:
assert parse_identifier(written) == ("isbn-13", "9780486282114")
def test_the_scheme_attribute_is_read_too(self) -> None:
assert parse_identifier("0-486-28211-2", "ISBN") == ("isbn-10", "0486282112")
def test_non_isbn_identifiers_are_kept(self) -> None:
assert parse_identifier("urn:uuid:3f2b1c4e-1111-2222-3333-444455556666") == (
"uuid",
"3f2b1c4e-1111-2222-3333-444455556666",
)
assert parse_identifier("calibre:1234") == ("calibre", "1234")
assert parse_identifier("B000FC0PDA", "mobi-asin") == ("asin", "B000FC0PDA")
def test_an_unrecognised_prefix_is_part_of_the_value(self) -> None:
""""http://example.com/book" is not an identifier called "http"."""
assert parse_identifier("http://www.gutenberg.org/5200") == (
"id",
"http://www.gutenberg.org/5200",
)
def test_a_declared_isbn_that_is_not_one_is_dropped(self) -> None:
assert parse_identifier("urn:isbn:not-an-isbn") is None
@pytest.mark.parametrize("written", ["", " ", None])
def test_nothing_yields_nothing(self, written: str | None) -> None:
assert parse_identifier(written) is None
+176 -1
View File
@@ -1,7 +1,13 @@
import pytest
from ebooklib import epub
from pathlib import Path
from datetime import date
from chitai.services.metadata_extractor import EpubExtractor
from chitai.services.metadata_extractor import (
EpubExtractor,
Extractor,
PdfExtractor,
split_edition,
)
@pytest.mark.asyncio()
@@ -15,3 +21,172 @@ class TestEpubExtractor:
assert metadata["authors"] == ["Herman Melville"]
assert metadata["published_date"] == date(year=2001, month=7, day=1)
EPUB = Path("tests/data_files/Metamorphosis - Franz Kafka.epub")
PDF = Path("tests/data_files/Calculus Made Easy - Silvanus Thompson.pdf")
@pytest.mark.asyncio()
class TestIdentifierMerging:
"""A book's formats each contribute identifiers; none of them replaces the rest."""
async def test_every_format_contributes(self, monkeypatch: pytest.MonkeyPatch) -> None:
"""
Identifiers are a collection, not a single value.
Merging the whole dict meant the last format to report won outright: an EPUB
declaring an ASIN, a Google volume id and a Calibre id kept none of them once
a PDF contributed one ISBN.
"""
async def epub(_file):
return {
"title": "How Linux Works",
"identifiers": {"isbn-13": "9781718500419", "asin": "1718500408"},
}
async def pdf(_file):
return {"identifiers": {"isbn-13": "9781718500402", "isbn-10": "1593270356"}}
monkeypatch.setattr(EpubExtractor, "extract_metadata", epub)
monkeypatch.setattr(PdfExtractor, "extract_metadata", pdf)
metadata = await Extractor.extract_metadata(
[Path("How Linux Works.epub"), Path("How Linux Works.pdf")]
)
assert metadata["identifiers"] == {
# Declared by the publisher's toolchain, so it outranks the PDF's, which
# was scraped off a copyright page that also prints the print edition's.
"isbn-13": "9781718500419",
"asin": "1718500408",
"isbn-10": "1593270356",
}
async def test_a_second_format_that_finds_nothing_erases_nothing(self) -> None:
"""The PDF fixture carries no ISBN, so it must leave the EPUB's alone."""
metadata = await Extractor.extract_metadata([EPUB, PDF])
assert metadata["identifiers"] == {"id": "http://www.gutenberg.org/5200"}
async def test_one_format_on_its_own_is_unaffected(self) -> None:
metadata = await Extractor.extract_metadata([EPUB])
assert metadata["identifiers"] == {"id": "http://www.gutenberg.org/5200"}
async def test_no_identifiers_anywhere_leaves_the_field_absent(self) -> None:
"""An empty dict would count as extracted metadata and overwrite nothing."""
metadata = await Extractor.extract_metadata([PDF])
assert "identifiers" not in metadata
class TestSplitEdition:
"""An edition is a field on the book, not part of what the book is called."""
@pytest.mark.parametrize(
("title", "stripped", "edition"),
[
# Every form below is one that turned up in a real library.
("Fluent Python, 2nd Edition", "Fluent Python", 2),
("Building Microservices, 2E", "Building Microservices", 2),
("Digital Image Processing, 4e", "Digital Image Processing", 4),
(
"Network Security Essentials: Applications and Standards/6e",
"Network Security Essentials: Applications and Standards",
6,
),
(
"Refactoring: Improving the Design of Existing Code (2nd edition)",
"Refactoring: Improving the Design of Existing Code",
2,
),
# Ordinal words, including a qualifier sitting inside the statement.
(
"The Art of Computer Programming: Volume 1 / Fundamental Algorithms, Third Edition",
"The Art of Computer Programming: Volume 1 / Fundamental Algorithms",
3,
),
(
"Introduction to the Theory of Computation, Third International Edition",
"Introduction to the Theory of Computation",
3,
),
# Mid-title, before a subtitle and before a trailing author.
(
"How Linux Works, 3rd Edition: What Every Superuser Should Know",
"How Linux Works: What Every Superuser Should Know",
3,
),
(
"Code Complete, 2nd Edition - Steve McConnell",
"Code Complete - Steve McConnell",
2,
),
# An underscore between the number and the "e", beside an unnumbered
# qualifier that has nowhere to go in an integer column and so stays put.
(
"Cryptography and Network Security, Global Edition, 8_e - Stallings",
"Cryptography and Network Security, Global Edition - Stallings",
8,
),
],
)
def test_editions_are_split_out(self, title: str, stripped: str, edition: int) -> None:
assert split_edition(title) == (stripped, edition)
@pytest.mark.parametrize(
"title",
[
# A number alone is never an edition — these are titles.
"Catch 22",
"Fahrenheit 451",
"Blade Runner 2049",
"1984",
"Apollo 13",
"Slaughterhouse 5",
"The Art of Computer Programming: Volume 1",
# "Edition" with no number cannot be stored, so it stays where it can
# still be read.
"Cryptography and Network Security: Principles and Practice, Global Edition",
"Building Microservices",
],
)
def test_titles_are_left_alone(self, title: str) -> None:
assert split_edition(title) == (title, None)
def test_a_title_that_is_only_an_edition_is_kept(self) -> None:
"""Stripping must never leave a book with no title at all."""
assert split_edition("2nd Edition") == ("2nd Edition", None)
@pytest.mark.parametrize("title", ["", None])
def test_nothing_yields_nothing(self, title: str | None) -> None:
assert split_edition(title) == (title, None)
@pytest.mark.asyncio()
class TestEditionFromFiles:
async def test_extraction_moves_the_edition_off_the_title(self) -> None:
"""The PDF fixture calls itself a 2nd edition in its own metadata title."""
metadata = await Extractor.extract_metadata([PDF])
assert metadata["title"] == "The Project Gutenberg eBook #33283: Calculus Made Easy"
assert metadata["edition"] == 2
class TestEpubPublisher:
"""The publisher was looked up and then dropped on the floor."""
def test_a_declared_publisher_is_returned(self) -> None:
"""
The lookup discarded its own result and fell off the end of the function, so
every EPUB reported no publisher no matter what it said.
"""
book = epub.EpubBook()
book.add_metadata("DC", "publisher", "No Starch Press")
assert EpubExtractor._extract_publisher(book) == "No Starch Press"
def test_no_publisher_is_none(self) -> None:
assert EpubExtractor._extract_publisher(epub.EpubBook()) is None
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,415 @@
"""Tests for importing a Calibre library through BookService."""
from pathlib import Path
import pytest
from chitai.config import settings
from chitai.database import models as m
from chitai.services import BookService
from chitai.services.calibre import CalibreLibrary
from tests.calibre_fixtures import CalibreFixture
DATA_FILES = Path("tests/data_files")
EPUB = DATA_FILES / "Metamorphosis - Franz Kafka.epub"
OTHER_EPUB = DATA_FILES / "The Art of War - Sun Tzu.epub"
PDF = DATA_FILES / "Calculus Made Easy - Silvanus Thompson.pdf"
@pytest.fixture(name="calibre_root")
def fx_calibre_root(tmp_path: Path) -> Path:
"""Three books: one plain, one in two formats, one Chitai cannot use."""
fixture = CalibreFixture(tmp_path / "source")
fixture.add_pages_table()
fixture.add_book(
1,
"The Metamorphosis",
authors=["Franz Kafka"],
pubdate="1915-10-15 00:00:00+00:00",
tags=["Fiction", "Absurdist"],
publisher="Kurt Wolff Verlag",
languages=["deu"],
comment="<p>He wakes up <i>changed</i>.</p>",
identifiers={"isbn": "978-0-486-29030-0", "amazon": "B01N5IB20Q"},
uuid="11111111-2222-3333-4444-555555555555",
pages=201,
cover=True,
formats={"EPUB": EPUB},
)
fixture.add_book(
2,
"The Art of War",
authors=["Sun Tzu"],
series="Classics",
series_index=3.0,
formats={"EPUB": OTHER_EPUB, "PDF": PDF},
)
fixture.add_book(3, "Metadata Only", authors=["Nobody"])
return fixture.commit()
async def library_of(root: Path) -> CalibreLibrary:
source = CalibreLibrary(root)
await source.open()
return source
async def test_imports_a_catalogue(
books_service: BookService, test_library: m.Library, calibre_root: Path
) -> None:
source = await library_of(calibre_root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
assert result.total == 3
assert len(result.created) == 2
# The book with no files is left out: a record with nothing to read, and a directory
# to match, is worse than not importing it.
assert [skipped.calibre_id for skipped in result.skipped] == [3]
assert result.skipped[0].reason == "no files in the catalogue"
assert result.failed == []
book = await books_service.get(result.created[0])
assert book.title == "The Metamorphosis"
assert [author.name for author in book.authors] == ["Franz Kafka"]
assert sorted(tag.name for tag in book.tags) == ["Absurdist", "Fiction"]
assert book.publisher is not None and book.publisher.name == "Kurt Wolff Verlag"
assert book.published_date is not None and book.published_date.year == 1915
assert book.language == "deu"
assert book.pages == 201
assert book.cover_image is not None
# The HTML is gone; `Book.description` is rendered as text.
assert book.description == "He wakes up changed."
async def test_two_formats_are_one_book(
books_service: BookService, test_library: m.Library, calibre_root: Path
) -> None:
source = await library_of(calibre_root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
book = await books_service.get(result.created[1])
assert book.title == "The Art of War"
assert sorted(Path(file.path).suffix for file in book.files) == [".epub", ".pdf"]
# A REAL series index reaches the column as the string everything else writes.
assert book.series is not None and book.series.title == "Classics"
assert book.series_position == "3"
async def test_identifiers_are_folded_onto_chitai_schemes(
books_service: BookService, test_library: m.Library, calibre_root: Path
) -> None:
"""
`amazon` becomes `asin`, a hyphenated ISBN survives, and the Calibre uuid is kept.
The uuid is deliberately not stored under `uuid`, which duplicate matching ignores
because an EPUB regenerates one per build. Calibre's is stable, so it is the durable
link back to the row it came from.
"""
source = await library_of(calibre_root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
book = await books_service.get(result.created[0])
identifiers = {identifier.name: identifier.value for identifier in book.identifiers}
assert identifiers["asin"] == "B01N5IB20Q"
assert identifiers["isbn-13"] == "9780486290300"
assert identifiers["calibre-uuid"] == "11111111-2222-3333-4444-555555555555"
matching = {
identifier.name: identifier.normalized_value for identifier in book.identifiers
}
# Stored under its own name, matched under one scheme for both ISBN forms.
assert matching["isbn-13"] == "isbn:9780486290300"
# And the uuid carries a real matching key, which is the whole reason it is not
# filed under `uuid`.
assert matching["calibre-uuid"] is not None
async def test_the_source_library_is_left_alone(
books_service: BookService, test_library: m.Library, calibre_root: Path
) -> None:
"""Files are copied. Moving them would leave `metadata.db` pointing at nothing."""
before = {
path: path.stat().st_mtime_ns
for path in sorted(calibre_root.rglob("*"))
if path.is_file()
}
source = await library_of(calibre_root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
after = {
path: path.stat().st_mtime_ns
for path in sorted(calibre_root.rglob("*"))
if path.is_file()
}
assert after == before
# And the copies are really there, under the library's own layout.
for book_id in result.created:
book = await books_service.get(book_id)
for file in book.files:
assert (Path(book.path or "") / file.path).is_file()
async def test_importing_twice_creates_nothing(
books_service: BookService, test_library: m.Library, calibre_root: Path
) -> None:
"""
Re-running is safe with no bookkeeping: the bytes are recognised wherever they sit.
This is what makes an interrupted import resumable by simply running it again.
"""
for _ in range(2):
source = await library_of(calibre_root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
assert result.created == []
assert sorted(skipped.reason for skipped in result.skipped) == [
"already stored",
"already stored",
"no files in the catalogue",
]
held_by = [
skipped.book_id for skipped in result.skipped if skipped.reason == "already stored"
]
assert all(book_id is not None for book_id in held_by)
async def test_a_file_the_catalogue_lists_but_disk_does_not(
books_service: BookService, test_library: m.Library, tmp_path: Path
) -> None:
"""Calibre keeps the row when a file is moved away behind its back."""
fixture = CalibreFixture(tmp_path / "source")
fixture.add_book(1, "Present", authors=["A"], formats={"EPUB": EPUB})
fixture.add_book(2, "Absent", authors=["B"])
fixture.add_missing_format(2, "EPUB", "Absent - B")
root = fixture.commit()
source = await library_of(root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
assert len(result.created) == 1
assert [(s.calibre_id, s.reason) for s in result.skipped] == [(2, "no files on disk")]
async def test_one_broken_book_does_not_stop_the_import(
books_service: BookService, test_library: m.Library, tmp_path: Path
) -> None:
"""
A failure is recorded and the run continues, leaving no files behind for it.
An orphaned directory would make the next attempt reserve `title (2)` and look as
though it had worked.
"""
fixture = CalibreFixture(tmp_path / "source")
fixture.add_book(1, "First", authors=["A"], formats={"EPUB": EPUB})
fixture.add_book(2, "Doomed", authors=["B"], formats={"EPUB": OTHER_EPUB})
fixture.add_book(3, "Third", authors=["C"], formats={"PDF": PDF})
root = fixture.commit()
original = books_service.create
async def fail_on_the_second(data, *args, **kwargs):
if isinstance(data, dict) and data.get("title") == "Doomed":
raise RuntimeError("no room on the shelf")
return await original(data, *args, **kwargs)
books_service.create = fail_on_the_second # type: ignore[method-assign]
source = await library_of(root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
books_service.create = original # type: ignore[method-assign]
assert len(result.created) == 2
assert len(result.failed) == 1
assert result.failed[0].calibre_id == 2
assert "no room on the shelf" in result.failed[0].reason
# Nothing of the failed book was left in the library. Checked against the path the
# template would have produced, rather than by walking the root — the Calibre source
# sits under it in these tests, and its own files are meant to still be there.
assert not (Path(test_library.root_path) / "B").exists()
# And the books either side of it are where they should be.
for book_id in result.created:
book = await books_service.get(book_id)
for file in book.files:
assert (Path(book.path or "") / file.path).is_file()
async def test_a_cover_that_cannot_be_read_is_not_fatal(
books_service: BookService, test_library: m.Library, tmp_path: Path
) -> None:
"""
A truncated `cover.jpg` costs the cover, not the book.
Real libraries hold them, from an interrupted download or a failed conversion, and
the cover is the one thing in the directory that can be replaced from the book page.
"""
fixture = CalibreFixture(tmp_path / "source")
fixture.add_book(
1, "Unreadable Cover", authors=["A"], corrupt_cover=True, formats={"EPUB": EPUB}
)
root = fixture.commit()
source = await library_of(root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
assert result.failed == []
assert len(result.created) == 1
book = await books_service.get(result.created[0])
assert book.cover_image is None
assert len(book.files) == 1
async def test_shared_authors_and_tags_are_one_row_each(
books_service: BookService, test_library: m.Library, tmp_path: Path, session
) -> None:
"""Two books by one author must not produce two `Author` rows."""
fixture = CalibreFixture(tmp_path / "source")
fixture.add_book(
1, "One", authors=["Franz Kafka"], tags=["Fiction"], formats={"EPUB": EPUB}
)
fixture.add_book(
2, "Two", authors=["Franz Kafka"], tags=["Fiction"], formats={"EPUB": OTHER_EPUB}
)
root = fixture.commit()
source = await library_of(root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
assert len(result.created) == 2
first, second = [await books_service.get(book_id) for book_id in result.created]
assert first.authors[0].id == second.authors[0].id
assert first.tags[0].id == second.tags[0].id
async def test_a_second_copy_is_reported_not_refused(
books_service: BookService, test_library: m.Library, tmp_path: Path
) -> None:
"""
Two catalogue rows for one book, with different bytes, both import.
File-level dedupe cannot see it the archives differ so book-level detection
reports the pair and leaves the decision to the reader.
"""
padded = tmp_path / "padded.epub"
padded.write_bytes(EPUB.read_bytes() + b"\0" * 64)
fixture = CalibreFixture(tmp_path / "source")
fixture.add_book(
1, "The Metamorphosis", authors=["Franz Kafka"], formats={"EPUB": EPUB}
)
fixture.add_book(
2, "The Metamorphosis", authors=["Franz Kafka"], formats={"EPUB": padded}
)
root = fixture.commit()
source = await library_of(root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
assert len(result.created) == 2
assert len(result.possible_duplicates) == 1
assert result.possible_duplicates[0].candidates[0].book_id == result.created[0]
async def test_allow_duplicates_stores_the_same_bytes_again(
books_service: BookService, test_library: m.Library, calibre_root: Path
) -> None:
for allow in (False, True):
source = await library_of(calibre_root)
try:
result = await books_service.create_many_from_calibre(
source, test_library, allow_duplicates=allow
)
finally:
await source.close()
assert len(result.created) == 2
async def test_duplicate_scope_off_imports_everything(
books_service: BookService,
test_library: m.Library,
calibre_root: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
monkeypatch.setattr(settings, "duplicate_scope", "off")
for _ in range(2):
source = await library_of(calibre_root)
try:
result = await books_service.create_many_from_calibre(source, test_library)
finally:
await source.close()
assert len(result.created) == 2
async def test_progress_is_reported_per_book(
books_service: BookService, test_library: m.Library, calibre_root: Path
) -> None:
"""The import is long enough that its progress is the only thing worth watching."""
seen = []
source = await library_of(calibre_root)
try:
await books_service.create_many_from_calibre(
source, test_library, on_progress=seen.append
)
finally:
await source.close()
assert [progress.processed for progress in seen] == [1, 2, 3]
assert all(progress.total == 3 for progress in seen)
assert [progress.outcome for progress in seen] == ["created", "created", "skipped"]
assert seen[0].title == "The Metamorphosis"
+506 -448
View File
File diff suppressed because it is too large Load Diff
+407
View File
@@ -0,0 +1,407 @@
# Implementation brief: importing a Calibre library
Written for an agent picking this up cold. Read the repo-root `AGENTS.md` and
`backend/AGENTS.md` first — this brief assumes both, particularly the **Filesystem behaviour**
and **Duplicate detection** sections.
Every claim about Calibre's schema and on-disk layout below was checked against a real library at
`~/Documents/Calibre Library` (6 books, current Calibre). Where a fact came from Calibre's source
rather than that library, it says so.
## Feasibility: high, and most of the machinery already exists
`metadata.db` is plain SQLite with a schema that has been stable for a decade, and the files sit
beside it in a predictable tree. Chitai already has every piece an import needs:
| Needed | Already in the tree |
| --- | --- |
| Ingest files that are already on disk | `BookService.create_many_from_existing_files` (`services/book.py:1280`) |
| Deduplicate authors/tags/publishers/series | `_populate_with_unique_relationships` (`services/book.py:1752`), via `as_unique_async` |
| Generic identifiers with a scheme map | `Identifier`, and `parse_identifier` (`services/metadata_extractor.py:58`) — whose map already covers `isbn`, `amazon`, `mobi-asin`, `google`, `goodreads`, `doi`, `calibre` |
| Decide where a book lives on disk | `BookPathGenerator`, `_reserve_book_path` (`services/book.py:1066`) |
| Not import the same book twice | `find_duplicate_files` (`:461`) and `find_duplicate_books` (`:535`) |
| Store a cover | `_save_cover_image` (`:1994`) |
So this is **a reader, not new ingest machinery**: turn Calibre rows into the metadata dict
`BookService` already accepts, and hand it to a slightly generalised version of the consume-directory
path. The parsing is the easy half.
The hard parts are elsewhere, and all three are addressed below:
1. Calibre libraries hold formats Chitai cannot describe, let alone read — and one of them
**currently 500s the book detail endpoint** (see prerequisites).
2. A 5,000-book import is a long-running job, and the app has no job/progress concept.
3. Whether files are **copied** into the library or **referenced in place** — which is a
product decision with a large blast radius, because in-place means Chitai's write paths point
at somebody's Calibre library.
## What a Calibre library actually is
```
Calibre Library/
├── metadata.db the whole catalogue
├── metadata_db_prefs_backup.json ignore
├── .caltrash/ .calnotes/ ignore — deleted books still live in .caltrash
└── <Author Name>/
└── <Title> (<book id>)/ == books.path
├── cover.jpg iff books.has_cover
├── metadata.opf ignore; the db is authoritative
└── <data.name>.<format> one per row in `data`
```
The tables that matter, and nothing else: `books`, `authors` + `books_authors_link`,
`publishers` + `books_publishers_link`, `tags` + `books_tags_link`, `series` +
`books_series_link`, `languages` + `books_languages_link`, `comments`, `identifiers`, `data`,
`books_pages_link`, `last_read_positions`.
Ten things that will produce wrong data if you do not know them:
- **Never query the views.** `meta`, `tag_browser_*` and friends call SQLite functions Calibre
registers from Python at connection time. Verified: `SELECT * FROM meta` fails with
`no such function: sortconcat`. Query base tables only.
- **`pubdate` has a sentinel, not a null.** An unknown publication date is stored as
`0101-01-01 00:00:00+00:00` (Calibre's `UNDEFINED_DATE`, year 101). It parses fine as a
`date`, so nothing will complain — two of the six books in the reference library carry it. Drop
any `pubdate` with year < 1000. The same sentinel appears in `timestamp`.
- **`data.name` is lossy and is not the title.** It is the on-disk stem, truncated to Calibre's
filename limit and sanitised. Verified in the reference library: the book titled
`The Project Gutenberg eBook #33283: Calculus Made Easy, 2nd Edition` is stored as
`The Project Gutenberg eBook #33283_ Calcul - Silvanus Phillips Thompson.pdf`. So the file
extractors must not be consulted for metadata (see decision 2), and `books.path`/`data.name`
are for *locating* files only.
- **`books.title` may contain characters Calibre strips from its own paths** — `:` became `_`
above, and titles legitimately contain `/` (`AC/DC`). `BookPathGenerator` interpolates the
title straight into a path and only collapses repeated slashes
(`services/filesystem_library.py`), so an unsanitised Calibre title can silently add a
directory level. Sanitise `/` and control characters out of `title` before path generation.
- **`authors.name` escapes commas as `|`.** Calibre's `AuthorsTable` unserialises with
`name.replace('|', ',')` (from Calibre's `db/tables.py`; the reference library has no such
name, so this one is unverified locally). Do the same replacement, and pass `authors.name`
**not** `authors.sort`, which is `Melville, Herman`. `format_author_name` would flip the sort
form correctly anyway, but there is no reason to hand it the worse input.
- **`series_index` is a REAL.** `7.0` must become `"7"`, not `"7.0"``Book.series_position` is
a string, and `find_duplicate_books`'s series-position disqualifier compares it as one.
- **`languages.lang_code` is ISO 639-2/B** (`eng`), while `EpubExtractor` stores raw
`DC:language` (`en`). Both will coexist in the column. `Book.language` is free text and the
edit form is a plain `<input>`, so nothing breaks; normalising to two letters is optional
polish, not part of this work.
- **`comments.text` is HTML.** `Book.description` is rendered as plain text by
`CollapsibleText`, so `<p>` tags will show literally. Strip to text on import.
- **`books_pages_link` is usually empty of real data.** It carries `needs_scan` and, in the
reference library, `pages = 0` for all six books. Only use it when `pages > 0`.
- **`identifiers.type` is free text.** The reference library holds `isbn`, `amazon` and
`mobi-asin`, all of which `parse_identifier` already maps. Feed every identifier through it and
keep whatever survives; do not filter to a known list.
## Decisions to settle before writing code
1. **Copy files into the library. Do not move, do not reference in place** — for the first
version. Moving leaves `metadata.db` pointing at files that are gone, which quietly destroys a
library the user still uses. Referencing in place is genuinely desirable (nobody wants two
copies of 80 GB) but it points `book.path` at the Calibre tree, and `update_book` **moves
directories** while `delete_books` **deletes files** — so a metadata edit in Chitai would
rearrange somebody's Calibre library. `Library.read_only` exists but is enforced in exactly one
place (`services/library.py:54`, at creation). See phase 3.
2. **Trust Calibre's metadata; do not run the extractors.** Calibre's catalogue is curated, its
filenames are truncated garbage, and running `Extractor.extract_metadata` over thousands of
files means opening every EPUB and rendering a cover page from every PDF. Take the cover from
`cover.jpg` directly. The one exception worth allowing: fill `pages` from the file when
Calibre has no useful value, behind a flag, off by default.
3. ~~**The source is a server-side path, not an upload.**~~ **Reversed in review, and the reason
this brief was wrong is worth keeping.** The premise — "the library lives on the same host as
the backend in every realistic deployment" — is false for the common case: Calibre is a desktop
application, and its library is on the desktop. So the split is by *surface*, not by preference:
- **Over HTTP: an uploaded zip only.** `…/imports/calibre/upload`. There is no endpoint taking a
server path; one was built and then removed deliberately.
- **On the server: a path only.** `litestar calibre-import <path>`, which is where a very large
library or a headless migration belongs — an upload has to carry the whole archive across
first.
A CLI taking a path needs no justification. An *endpoint* taking one would have: it would let any
authenticated caller read any directory the backend can, and `TODO.md` records there is no
authorization tier at all. Not adding it is one less thing to gate later.
4. **Import into an existing Chitai library**, chosen by the caller. Creating a library is
already one action, and the Calibre tree is rewritten by `BookPathGenerator` regardless.
5. **Re-running an import must be safe, and file-level dedupe already makes it so.** The same
bytes are recognised by `(hash, size)` whatever their path, so a second run over the same
library skips everything. No import bookkeeping is needed for idempotency.
## Prerequisite: a `.mobi` file breaks the book endpoint — **done**
> Landed ahead of the import itself. `guess_content_type` in `services/utils.py` now names every
> format from its extension, the column keeps a null when nothing can name one, and the OPDS feed
> substitutes `application/octet-stream` at the one place a string is required. The Read control is
> driven by `isReadable` rather than by the file count. See the section below for why it mattered,
> and `backend/AGENTS.md` for the rule as it now stands.
`FileMetadataRead.content_type` is a required `str` (`schemas/book.py:33`), but every ingest path
fills it from `mimetypes.guess_type`, which returns `None` for `.mobi`, `.azw`, `.fb2`, `.lit`
and `.htmlz` (verified). `FileMetadata.content_type` is nullable in the model, so the row stores
fine and then fails response validation on the way out — a book whose only file is a MOBI would
be unreadable through the API.
Nothing in the tree hits this today because the browser upload path is used with EPUBs and PDFs.
A Calibre library is full of MOBI and AZW3. Fix it first, either way round:
- make the schema field `str | None`, and/or
- add a small extension→MIME table for the ebook formats `mimetypes` does not know
(`application/x-mobipocket-ebook`, `application/vnd.amazon.ebook`, `application/x-fictionbook+xml`).
Do both, in fact: the table is the right answer for OPDS clients, which choose an acquisition link
by MIME type, and the nullable field is the safety net.
**Related, but not a blocker:** Chitai reads EPUB and PDF only. `openBookInReader`
(`book/[bookId]/+page.svelte:66`) branches on `getFileType(...) === 'EPUB' | 'PDF'` and does
nothing for anything else, so an AZW3-only book gets a Read button that silently fails. Importing
those files is still right — they are downloadable and they are the user's — but the button
should be disabled for a book with no readable file. One `$derived` on the page, worth doing in
the same branch.
## Design
### 1. `services/calibre.py` — a pure reader, no Chitai types
```python
@dataclass(frozen=True)
class CalibreFile:
path: Path # absolute, resolved against the library root
format: str # "EPUB", as stored
size: int # data.uncompressed_size, for a cheap sanity check
@dataclass(frozen=True)
class CalibreBook:
calibre_id: int
uuid: str
title: str
authors: list[str]
... # one field per row of the mapping table below
cover: Path | None
files: list[CalibreFile]
class CalibreLibrary:
def __init__(self, root: Path) -> None: ...
async def open(self) -> None: ... # copy + connect, see below
async def books(self) -> AsyncIterator[CalibreBook]: ...
async def close(self) -> None: ...
```
Deliberately knows nothing about `Book`, `BookService` or the session — it is a file-format
reader, unit-testable against a fixture database with no Postgres and no app.
Two implementation notes:
- **Copy `metadata.db` to a temp file and read the copy.** Calibre may be running and writing;
opening the live file read-only either sees a torn state or needs the `-wal` sidecar. The
database is small (438 KB for six books, single-digit MB for thousands), so a copy costs
nothing and removes the whole problem.
- **`sqlite3` inside `asyncio.to_thread`, not a new dependency.** The connection is used for a
handful of queries. Do not add `aiosqlite` for this.
Read the whole catalogue in **one query per table** and join in Python — six or so `SELECT`s and
a few dicts, versus a per-book N+1 across ten tables. At self-hosted scale the entire catalogue
minus descriptions fits in memory comfortably; if `comments.text` for 20k books is a concern,
fetch that one table per batch.
### 2. The mapping
| Calibre | Chitai | Notes |
| --- | --- | --- |
| `books.title` | `title` | Sanitise `/` and control chars for path generation. `Extractor.format_book_title` may still be worth applying to split a subtitle at the second colon — but **do not** run `split_edition`, Calibre's title is the curated one. |
| `authors.name` via `books_authors_link` | `authors` | `\|``,`. Order by `books_authors_link.id`; `Book.author_links` is an `ordering_list`, so insertion order is the displayed order. |
| `comments.text` | `description` | Strip HTML to text. |
| `books.pubdate` | `published_date` | Drop the year-101 sentinel. |
| `series.name`, `books.series_index` | `series`, `series_position` | `7.0``"7"`. |
| `tags.name` | `tags` | |
| `publishers.name` | `publisher` | `books_publishers_link` is unique per book. |
| `languages.lang_code` (lowest `item_order`) | `language` | Chitai holds one. |
| `identifiers.type` / `.val` | `identifiers` | Through `parse_identifier`; keep what survives. |
| `books.uuid` | `identifiers["calibre-uuid"]` | The one durable link back to the source row. `normalize_identifier` returns a key for it (it is not in `_PER_BUILD_NAMES`), which is *desirable*: a book re-imported from the same Calibre library matches on it exactly. |
| `books_pages_link.pages` | `pages` | Only when `> 0`. |
| `cover.jpg` when `has_cover` | `cover_image` | |
| `data` rows | `files` | |
| `books.timestamp` | — | `Book.created_at` is audit-managed; do not fight it. |
| `ratings`, `annotations`, `custom_columns` | — | No home in the model. Out of scope. |
| `last_read_positions` | `BookProgress` | Phase 2. |
### 3. The ingest
> **As built, this went on `BookService` as `create_many_from_calibre`, not into a separate
> `services/calibre_import.py`.** The orchestration needs `_reserve_book_path`,
> `_save_cover_image`, `_screen_for_duplicates` and `_record_possible_duplicates`, and reaching
> into four privates from another module is worse than one more method in the file where the other
> two ingest paths already live. `_record_possible_duplicates` was changed to take the list it
> appends to rather than an `ImportResult`, so every ingest path can share it whatever its own
> result type is. Two other deviations: `CalibreLibrary.books()` returns a list rather than an
> async iterator, because the caller needs the total up front anyway; and the reader reports
> identifiers exactly as Calibre keyed them, with the fold onto Chitai's schemes done by the
> importer, which keeps the reader free of Chitai imports.
Per book, in this order — it mirrors `create_many_from_existing_files`, which is the closest
existing shape:
1. `fingerprint_file` each source file (`services/utils.py:164`).
2. `find_duplicate_files` against the target library. All files known → skip the book entirely,
recording it. Some known → import the rest.
3. Build the metadata dict from the `CalibreBook`.
4. `_reserve_book_path(path_gen.generate_path(data))`.
5. **Copy** each file to `parent / _unused_path(...)`, building `FileMetadata` from the
fingerprint already computed. `services/utils.py` has `move_file` but no copy — add
`copy_file` beside it, streaming through `aiofiles` in `CHUNK_SIZE` blocks like
`_save_book_files` does, not `shutil.copy` (a 40 MB blocking read inside the event loop).
6. Cover: open `cover.jpg` with PIL and hand the `Image` to `_save_cover_image`, which already
accepts one and converts to WebP.
7. `super().create(data)` through `BookService`, then `find_duplicate_books` and record
candidates — same as `_record_possible_duplicates` (`services/book.py:1253`).
8. **Commit per book.** A 5,000-book import inside one transaction is one failure away from
nothing, and per-book commits are what lets the library page show books arriving — which is
the behaviour commit `85367da` deliberately built.
Report an `ImportResult`-shaped outcome; reuse `ImportResult` itself if it fits, extending it with
a `failures: list[tuple[int, str]]` keyed by Calibre id. **One book must never fail the run**
a missing file, an unreadable cover or a `NOT NULL` violation gets recorded and skipped.
### 4. Progress, and where the import runs
> **As built, the HTTP surface is an uploaded archive and nothing else** — see decision 3.
>
> `POST …/imports/calibre/upload` takes a zipped library. The job owns the temp directory it is
> unpacked into and deletes it when it ends. Extraction refuses zip slip, an archive too big for the
> disk, and one with no `metadata.db` within three levels — all answered 400 before a job exists.
>
> This forced a fix to the SvelteKit proxy, which buffered request bodies with `arrayBuffer()`:
> survivable for one book, not for a multi-gigabyte archive. POST and PATCH now stream
> `request.body` through with `duplex: 'half'`.
>
> **A preview endpoint was built and then removed with the path route.** It read a server-side
> catalogue and reported its size before writing anything, which is only useful when the caller
> named a directory. An upload has already been carried across by the time anything can be read, so
> unpacking it *is* the validation step — an archive that is not a Calibre library is refused there.
The import outlives its request, so the handler starts it and returns a handle:
- `POST /libraries/{library_id:int}/imports/calibre` — body `{path, copy_files: true}`, returns
`{job_id, total}`. 202.
- `GET /libraries/imports/{job_id}` — `{state, total, processed, created, skipped, failed,
current_title, errors}`.
- `DELETE /libraries/imports/{job_id}` — cancel; the task checks a flag between books.
Keep the registry **in memory**, a `dict[str, ImportJob]` on a module-level singleton, with the
task created by `asyncio.create_task`. This matches what the app already does — the consume
watcher is an in-process singleton started from a lifespan hook — and it is roughly thirty lines
against a model, a migration and a service for the alternative.
State that limitation explicitly in the docstring: **it assumes one worker process.** The
production `CMD` is `litestar run`, which is single-process, so this holds today; `TODO.md`
already records that the production image should move to uvicorn with a worker count, and doing
that would mean a poll landing on a worker that has never heard of the job. The consume watcher
has the same problem, so this is not a new constraint — but the next person to add workers needs
to find it written down. If import *history* is ever wanted, that is when an `import_jobs` table
earns its migration.
**Also add a CLI entry point.** A 200 GB library imported through a browser tab that must stay
open is a bad experience, and `pyproject.toml` already declares a `chitai` script. A Litestar CLI
command (`litestar --app-dir src/chitai/ calibre-import <path> --library <slug>`) is ~20 lines
over the same service and is the right tool for the initial migration, which is the case this
whole feature exists for. The endpoint is for people who would rather click.
### 5. Frontend
Model it on the duplicates screen, which is the closest precedent in shape and placement:
- Route `(root)/settings/libraries/[libraryId]/import` — beside
`settings/libraries/[libraryId]/duplicates`, reached from the library settings page.
- `getCalibreImport` (a `query`) and `cancelCalibreImport` (a `command`) in
`src/lib/api/calibre-import.remote.ts`, re-exported from `src/lib/api/index.ts`. **Starting an
import is not a remote function**: the archive goes to the backend through the proxy so the
browser streams straight through, where a remote function would put the whole thing through the
SvelteKit process first. The screen uses `XMLHttpRequest` for it, which is the only way to get
upload progress.
- A file input, an upload progress bar, then a progress bar polling `getCalibreImport` every second
or two, a running count, and the failures listed at the end with their Calibre ids.
- Finish with a link to the library's duplicates screen. An import into a non-empty library is
the single most likely way to produce duplicate books, and that screen already handles them.
Do **not** route this through the upload tray. The tray reports on a client-driven queue it owns
(`upload-queue.svelte.ts`); this is server-side work whose state survives a page reload, and
conflating the two would mean teaching the tray to poll.
Regenerate `src/lib/schema/openapi/schema.d.ts` against a backend running **your** branch —
a stale server silently writes a stale file.
## Testing
The fixture is the interesting part. Build a Calibre library in a `tmp_path` fixture rather than
committing a binary `metadata.db`: a helper that executes the subset of Calibre's `CREATE TABLE`
statements (they are in this document's shape, and in any real library's `sqlite_master`), inserts
a handful of books, and lays out `<Author>/<Title> (id)/` directories containing the existing
EPUB and PDF fixtures from `backend/tests/data_files/` plus a copy of `cover.jpg`. Generated
beats committed here because the tests need to assert on *odd* rows — the pubdate sentinel, a
`|` in an author name, a title with a colon — and those are clearer written in Python than hidden
in a blob.
- **Unit** (`tests/unit/test_calibre.py`) — the reader alone: field mapping; the year-101 pubdate
dropped; `series_index` 7.0 → `"7"`; `|` unescaped in an author name; HTML stripped from
`comments`; `pages = 0` ignored; identifiers passed through `parse_identifier`; a `data` row
whose file is missing from disk reported rather than raised; `.caltrash` never walked.
- **Service** (`tests/unit/test_services/test_calibre_import.py`) — a book with two formats lands
as one record with two files; the source files still exist afterwards; a second run over the
same library creates nothing; a library with one broken book imports the rest; authors and tags
shared between two books produce one `Author` / `Tag` row each; `possible_duplicates` reported
when the target library already holds the same book.
- **Integration** (`tests/integration/test_calibre_import.py`) — `POST` returns 202 with a job id,
polling reaches a terminal state, and the books are then listable through `GET /books`. A
`.mobi`-only book must come back from `GET /books/{id}` without a 500 — that is the
prerequisite's regression test.
`pytest` needs Docker (`pytest-databases`).
## Verification
```bash
nix-shell # postgres + migrations applied
cd backend
pytest tests/ # take your own baseline first
uv run litestar --app-dir src/chitai/ run --port 8001 # for the OpenAPI regeneration
cd ../frontend && pnpm check # baseline: 30 errors, 1 warning, 8 files
```
Neither `pnpm check` nor `pnpm lint` is clean on this repo — baseline before assuming an error is
yours.
End to end, against a real library (`~/Documents/Calibre Library` will do): import into an empty
Chitai library and confirm the six books arrive with their covers, authors, tags, series
positions and identifiers intact; that the AZW3 book is listed and downloadable; that the source
library is byte-for-byte untouched (`diff -r` a copy taken beforehand); and that re-running the
import creates nothing and reports six skipped books. Then import the same library into a
library that already holds one of those books by upload, and confirm it lands on the duplicates
screen rather than as a second copy.
## Phasing
| Phase | Scope |
| --- | --- |
| **0** ✅ | The `content_type` prerequisite, plus the Read button. **Done** — see below. |
| **1** ✅ | `services/calibre.py`, the ingest, the CLI command, copy-only. **Done** — a headless one-time migration works today. |
| **2** ◐ | The upload endpoint, the job registry and the settings screen — **done**. The HTTP surface is an uploaded zip only; see decision 3, which this reversed. `last_read_positions``BookProgress` is **not** done: it needs a Calibre-user → Chitai-user mapping, and the answer differs between the endpoint (which has a `current_user`) and the CLI (which has none). That decision is the next thing to make. |
| **3** | Reference-in-place import. Its real content is **enforcing `Library.read_only`** across `update_book`, `delete_books`, `add_files` and `remove_files` — which is a feature of its own and should not be smuggled in under an import. |
## Out of scope
Annotations and highlights (no model to put them in), custom columns, ratings, virtual libraries
and saved searches → bookshelves, format conversion, writing anything back to Calibre, and any
form of continuing two-way sync. This is a one-way migration.
+146
View File
@@ -0,0 +1,146 @@
# Implementation brief: move Duplicates into library settings
Written for an agent picking this up cold. Read the repo-root `AGENTS.md` and
`frontend/AGENTS.md` first — this brief assumes both.
**This is a frontend-only change.** The backend already scopes everything by library
(`GET /books/duplicate-books?library_id=`), so no endpoint, schema or migration is
involved.
## Where this starts from
The duplicates review screen exists and works. It currently lives at
```
frontend/src/routes/(root)/(library)/library/[libraryId]/duplicates/
+page.server.ts loads the groups, plus the full Book records merge needs
+page.svelte group cards, "Not duplicates", "Merge…"
```
and is reached from a **Duplicates entry in the main sidebar**
(`frontend/src/lib/components/layout/nav-main.svelte`), which is what this change
removes.
Settings today is a flat, entirely global four-item nav
(`frontend/src/routes/(root)/settings/+layout.svelte`): Account, Appearance, Libraries,
Devices. `settings/libraries/+page.svelte` is a single table of every library whose rows
link *out* to the library itself. **There is nowhere that means "settings for this
library"** — that is the gap this change fills.
## What to build — option B
Libraries expands in the settings nav. Every library is a sub-item; selecting one swaps
the pane; Duplicates is a section inside that pane. All libraries and all their sections
end up one click apart.
```
/settings/libraries the existing table (leave it as the index)
/settings/libraries/[libraryId] redirects to the first section
/settings/libraries/[libraryId]/duplicates the review screen, moved
```
Suggested files:
| Path | What |
| --- | --- |
| `settings/libraries/[libraryId]/+layout.svelte` | Library name, and the section tabs |
| `settings/libraries/[libraryId]/+page.ts` | `redirect(303, …/duplicates)` |
| `settings/libraries/[libraryId]/duplicates/+page.server.ts` | Moved verbatim |
| `settings/libraries/[libraryId]/duplicates/+page.svelte` | Moved verbatim |
Duplicates is the **only** real section today. Build the tab strip so General and Danger
zone have somewhere obvious to land, but do not invent them now — an empty tab is worse
than no tab.
## The nav
In `settings/+layout.svelte`, `items` is a flat `as const` array matched on
`page.route.id`. Libraries needs to render its children beneath it:
```svelte
{#each libraryState.libraries as library (library.id)}
<a href={resolve('/(root)/settings/libraries/[libraryId]/duplicates', {
libraryId: String(library.id) })}> … </a>
{/each}
```
`getLibraryState()` **is** available under `/settings` — it is set in
`(root)/+layout.svelte`, above the settings group, and `settings/libraries/+page.svelte`
already uses it. No new load function is needed to list the libraries.
**Active state is matched on route id, not pathname.** There is a comment in
`settings/+layout.svelte` explaining why: `resolve()` returns an absolute path on the
client and a relative one during SSR, so a pathname comparison is false on the server and
true after hydration, and the highlight flashes in. A nested library item is active when
the route id matches **and** `page.params.libraryId === String(library.id)` — both, or
every library lights up at once.
## Things that will bite
1. **Remove the sidebar entry in the same change.** `nav-main.svelte` gained a
`Duplicates` item and a `CopyCheck` import when the screen was built. Delete both, and
delete the old route directory. Doing the removal and the move together is the point —
split across two commits the screen is unreachable in between.
2. **Delete the old route, do not leave it.** Two live copies of a screen that both write
is how they drift.
3. **`setBookSelectionState` is not available under `/settings`.** It is set in
`(root)/(library)/+layout.svelte`, which the settings group is not inside. This is
fine — the duplicates page uses `BookImage` directly, not `book-thumbnail.svelte`, and
`MergeBooks` takes its `libraryId` as a prop. **Verify this stays true** if you touch
either component; a `getBookSelectionState()` under settings returns `undefined` and
fails at the first access, not at import.
4. **Keep `depends('app:duplicate-books')`.** Both the dismiss action and `MergeBooks`
call `invalidate('app:duplicate-books')` to make a resolved group leave the screen.
Drop it and the page silently stops refreshing. `MergeBooks` also invalidates
`app:books`, which is a no-op under settings and should stay that way.
5. **The settings shell is height-constrained.** `settings/+layout.svelte` is
`h-[calc(100vh-var(--header-height)-2rem)]` with `overflow-auto` on the content pane.
The review screen is a long list of cards — it must scroll *inside* that pane. Its
current `mx-auto max-w-5xl` wrapper will want revisiting.
6. **Three levels of nav is option B's known cost.** Nav → library → section, and the
pane is narrower than the full-width route the screen was designed against. The group
cards are `w-36` covers in a wrapping flex row, so they reflow, but check a group of
four at a narrow window before calling it done.
7. **`resolve()` must be a direct call in markup** for `svelte/no-navigation-without-resolve`.
Where `nav-main.svelte` computes a url through a variable it carries an
`eslint-disable-next-line`; prefer the direct call over inheriting that.
8. **The loader depends on the `?ids=` fix.** `+page.server.ts` fetches full `Book`
records with `listBooks({ ids, pageSize })` because the merge workbench needs
identifiers, description and publisher, which `DuplicateBookRead` does not carry.
advanced_alchemy's stock id filter types that parameter as `list[str]` regardless of
config, which made Postgres refuse `bigint = character varying`; the override lives in
`backend/src/chitai/services/dependencies.py` (`create_book_filter_dependencies`).
If `GET /books?ids=1&ids=2` 500s, that override is missing — do not work around it in
the loader.
## Out of scope
The **General** and **Danger zone** sections (rename, path template, read-only, consume
directory, delete), and any change to the merge workbench itself. The toolbar entry point
for merge — select 2+ books in the library view — is unrelated and stays where it is.
## Verification
```bash
cd frontend
pnpm check # baseline: 30 errors, 1 warning, 8 files — none of them yours
pnpm lint # not clean either; check the files you touched, not the tree
pnpm build
```
By hand, with a library that has a duplicate group:
- Settings → Libraries lists every library beneath it; clicking one opens its pane.
- Duplicates shows the same groups the old route did, and scrolls inside the settings pane.
- **Not duplicates** removes the group and it stays gone after a reload.
- **Merge…** opens the workbench, merges, and the group leaves the screen.
- The main sidebar no longer has a Duplicates entry, and
`/library/<id>/duplicates` no longer resolves.
- A library with no duplicates shows the empty state, not a blank pane.
+3
View File
@@ -11,3 +11,6 @@ coverage
# Miscellaneous
/static/
# Vendored third-party source, copied verbatim by scripts/vendor-foliate.sh
/src/lib/vendor/
+178
View File
@@ -0,0 +1,178 @@
# Chitai frontend
SvelteKit web app for the eBook library. See the repo-root `AGENTS.md` for the overall picture and
dev-environment setup.
**Stack:** SvelteKit 2 with `adapter-node` · Svelte 5 (runes) · Tailwind v4 · Zod v4 ·
vendored `foliate-js` (EPUB) · vendored `pdf.js` (PDF) · `mode-watcher` (dark mode) ·
`svelte-sonner` (toasts) · pnpm.
Two experimental flags are on in `svelte.config.js` and the codebase depends on both:
`kit.experimental.remoteFunctions` and `compilerOptions.experimental.async` (`await` in components).
Tailwind v4 has **no config file** — the theme, colour tokens (hex, not oklch) and
`@custom-variant dark` all live in `src/app.css`. `dark` is the _only_ custom variant defined, so
generated components that assume others — shadcn's slider ships `data-horizontal:` / `data-vertical:`
classes — silently produce no styles. Use the `data-[orientation=…]` form instead.
## Talking to the backend
There are two mechanisms; pick deliberately.
**1. Remote functions — the default.** `src/lib/api/*.remote.ts` export `query` / `command` / `form`
functions from `$app/server`. Each takes a Zod schema from `$lib/schema` as its validator and runs
on the server, reaching the API through `locals.api` (the `ApiClient` in `src/lib/server/api.ts`):
```ts
export const getBook = query(stringCoerce, async (id): Promise<Book> => {
const { locals } = getRequestEvent();
const response = await locals.api.get(`/books/${id}`);
if (!response.ok) error(response.status === 404 ? 404 : 500, '…');
return await response.json();
});
```
Conventions: build query strings with `createQueryParams` from `$lib/utils`; on a failed response
throw SvelteKit's `error(status, message)`; multipart uploads go through `postMultipart` /
`putMultipart`. Re-export new modules from `src/lib/api/index.ts`.
**2. The catch-all proxy** at `src/routes/api/[...path]/+server.ts` forwards GET/POST/PATCH/DELETE
to the backend with the auth header attached. Use it only where the **browser itself** must fetch
the backend — e.g. streaming a book file into the EPUB/PDF reader. It is not the general-purpose
path.
## Auth
`src/hooks.server.ts` is a `sequence` of two handles: the first reads the `authToken` cookie,
constructs an `ApiClient`, validates it with `GET /access/me` and fills `locals.user` /
`locals.authToken` / `locals.api` (clearing the cookie if invalid); the second redirects any route
outside `/login` to the login page when there is no user.
The cookie is set in `src/lib/api/auth.remote.ts` (`login`) — httpOnly, secure, sameSite strict, one
week — and deleted by `logout`. The backend JWT never reaches client-side JS.
## Schemas
- `src/lib/schema/*.ts` — hand-written Zod schemas, used as remote-function input validators and as
the source of the exported TS types. Mirror `backend/src/chitai/schemas/` when the API changes.
- `src/lib/schema/common.ts` — shared building blocks: `stringCoerce` / `arrayCoerce` coercion
helpers, `PaginatedResponse<T>`, and the pagination / search / order query schemas that most list
endpoints compose from.
- `src/lib/schema/openapi/schema.d.ts`**generated** from the backend's OpenAPI document with
`openapi-typescript`. Never hand-edit; regenerate after backend API changes.
## State
Client state lives in classes in `src/lib/state/*.svelte.ts` using `$state` / `$derived`, shared via
Svelte context with a module-level `Symbol` key and a `setXState` / `getXState` pair:
```ts
const LIBRARY_KEY = Symbol('LIBRARY');
export function setLibraryState(libraries: Library[]) {
return setContext(LIBRARY_KEY, new LibraryState(libraries));
}
export function getLibraryState() {
return getContext<ReturnType<typeof setLibraryState>>(LIBRARY_KEY);
}
```
Follow that pattern rather than introducing stores. `library.svelte.ts` is the reference — including
its optimistic-delete-with-rollback and toast handling. `bookCollection` / `bookSelection` /
`bookOperations` split list data, selection and mutations across three cooperating classes.
## Components
- `src/lib/components/ui/` — vendored shadcn-svelte (`components.json`) plus jsrepo blocks from
`@ieedan/shadcn-svelte-extras` (`jsrepo.json`). Treat as generated: add components with the CLIs
rather than hand-writing them, and prefer wrapping over editing.
- App components live in `forms/`, `layout/`, `view/` (browser, grid/list/table, filters, sort) and
`reader/` (see [The readers](#the-readers)).
- `cn()` from `$lib/utils` merges Tailwind classes; the `WithElementRef` / `WithoutChild` helpers
there are the shadcn prop-typing conventions.
## The readers
**PDF** is the pdf.js viewer vendored under `static/pdfjs/`, pointed at by an iframe. Untouched by
the EPUB work; leave it alone unless the task is about PDFs.
**EPUB** is built on `foliate-js`, copied verbatim into `src/lib/vendor/foliate-js/` by
`scripts/vendor-foliate.sh` (pinned commit; see `src/lib/vendor/foliate-js/README.chitai.md`).
Upstream has no npm release and recommends a submodule; this repo has none and already vendors
pdf.js the same way, so it is copied instead. Only the import closure reachable from `view.js` is
vendored, and **`pdf.js` in that directory is our stub, not upstream's** — the real one imports a
bare `@pdfjs/pdf.min.mjs` that Rollup resolves at build time even though the path never runs.
Layout:
| Path | What |
| ------------------------------------------ | ---------------------------------------------------------------------------- |
| `lib/vendor/foliate-js/` | The engine. Do not edit — `vendor-foliate.sh` overwrites it. |
| `lib/reader/foliate.ts` | Lazy loader for the custom elements. The only thing that imports `$foliate`. |
| `lib/reader/settings.ts` · `stylesheet.ts` | Defaults/bounds, and the CSS injected into the book. |
| `lib/reader/progress.ts` | Debounced progress writer with a `sendBeacon` flush. |
| `lib/state/reader-settings.svelte.ts` | Settings state, persisted to `localStorage`. |
| `components/reader/foliate-view.svelte` | Wraps `<foliate-view>`; owns the imperative lifecycle. |
| `components/reader/epub-reader.svelte` | The shell: chrome, TOC, errors, progress. |
Things that will bite:
- **`$foliate` is a Vite-only alias.** It is deliberately absent from `kit.alias` and tsconfig
`paths` so TypeScript cannot resolve it and falls back to the ambient declaration in
`lib/reader/foliate-js.d.ts`; `src/lib/vendor` is also in tsconfig `exclude`. Without both,
`checkJs` walks ~11k lines of untyped JS. The declaration file must **not** be named `foliate.d.ts`
— beside `foliate.ts`, TypeScript takes it for that file's emitted declaration and drops it.
- **Never import the vendored code at module scope.** `view.js` calls `customElements.define` and
subclasses `HTMLElement` on import, so it must stay behind `loadFoliate()` inside `onMount`. SSR is
otherwise on for the reader route.
- **Sections render in iframes, which swallow key events.** Keyboard handlers are bound per section
document on the `load` event, and modifier combinations are replayed onto the host window so app
shortcuts (the sidebar's ctrl+B) still work while reading.
- **Renderer settings split two ways.** Flow, gap, margins, column count and line width are
_attributes_ set with `setAttribute` (there is no JS property API, no `margin` shorthand and no
`spread` — a spread is `max-column-count: 2`). Typography is CSS passed to `renderer.setStyles`,
which takes a `[before, after]` pair: the first is prepended to the section head so the book
overrides it, the second appended so it wins. User settings belong in the second, with
`!important`, or the book's own CSS beats them.
- **Progress needs no locations pre-pass.** `relocate` carries both a CFI and an overall `fraction`,
which map straight onto `epub_cfi` and `percentage`.
## Routing
Route groups carry the layout structure:
- `(root)` — the authenticated shell (sidebar, header); `+layout.server.ts` loads the libraries.
- `(root)/(library)` — library-scoped pages: `library/[libraryId]/view`, `book/[bookId]`, edit, and
the readers at `book/[bookId]/read/{epub,pdf}/[fileId]`.
- `+layout@.svelte` breakouts reset to the root layout for the login page and the full-screen reader.
## Conventions
Prettier (`.prettierrc`): tabs, single quotes, no trailing commas, 100 columns, with the Svelte and
Tailwind plugins. Run `pnpm check` (svelte-check) and `pnpm lint` before considering work done —
but take a baseline first, because neither is clean (see below).
`src/lib/vendor/` is excluded from Prettier, ESLint and svelte-check. Don't reformat vendored code.
## Known rough edges
Observed in the current tree — don't mistake these for intentional patterns to copy:
- `pnpm check` is not clean. **Baseline as of 2026-08-13: 30 errors, 1 warning, 8 files**, most of
them in `src/routes/api/[...path]/+server.ts` (see below). Get your own baseline before assuming
an error is yours. `src/lib/schema/openapi/schema.d.ts` was regenerated on that date and is
current; regenerate it again after any backend API change, with
`pnpm exec openapi-typescript http://localhost:8000/schema/openapi.json -o src/lib/schema/openapi/schema.d.ts`
against a backend running **your** branch — a stale server silently writes a stale file.
- `pnpm lint` does not pass either — 131 pre-existing ESLint errors, 35 of them
`svelte/no-navigation-without-resolve` on plain `href`s, plus ~59 files Prettier would rewrite
(mostly vendored shadcn components). Check the files you touched, not the whole tree.
- `src/routes/api/[...path]/+server.ts` — all four handlers are annotated `RequestHandler` while the
import of that type is commented out at line 4. It also buffers whole **responses** with
`arrayBuffer()` and forwards no `Range` header, so book downloads are not streamed. **Requests**
are streamed — POST and PATCH pass `request.body` through with `duplex: 'half'` (see `bodyOf`),
because a zipped Calibre library upload cannot be held in this process. The response side is
still buffered; see `TODO.md`.
- `src/app.d.ts``App.Locals["user"]` is typed from `lucide-svelte`'s `User` _icon_ component
rather than the `User` interface in `$lib/server/auth`.
- No CSP, which foliate's README asks for because EPUBs can carry scripts. See `TODO.md` for why it
is not enabled yet.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+8
View File
@@ -12,6 +12,8 @@ const gitignorePath = fileURLToPath(new URL('./.gitignore', import.meta.url));
export default defineConfig(
includeIgnoreFile(gitignorePath),
// Vendored third-party source. Tracked, so .gitignore does not cover it.
{ ignores: ['src/lib/vendor/**'] },
js.configs.recommended,
...ts.configs.recommended,
...svelte.configs.recommended,
@@ -27,6 +29,12 @@ export default defineConfig(
'no-undef': 'off'
}
},
{
// Generated shadcn components take href as a prop and cannot resolve it —
// that is the caller's job. Editing them here would be lost on regeneration.
files: ['src/lib/components/ui/**'],
rules: { 'svelte/no-navigation-without-resolve': 'off' }
},
{
files: ['**/*.svelte', '**/*.svelte.ts', '**/*.svelte.js'],
languageOptions: {
+3 -3
View File
@@ -33,8 +33,8 @@
"jsrepo": "^2.5.2",
"openapi-typescript": "^7.13.0",
"prettier": "^3.8.1",
"prettier-plugin-svelte": "^3.5.1",
"prettier-plugin-tailwindcss": "^0.6.14",
"prettier-plugin-svelte": "^3.5.2",
"prettier-plugin-tailwindcss": "^0.8.1",
"svelte": "^5.53.7",
"svelte-check": "^4.4.5",
"tailwind-merge": "^3.5.0",
@@ -47,7 +47,7 @@
"vite": "^7.3.1"
},
"dependencies": {
"epubjs": "^0.3.93",
"construct-style-sheets-polyfill": "^3.1.0",
"mode-watcher": "^1.1.0",
"svelte-sonner": "^1.0.8",
"zod": "^4.3.6"
+102 -291
View File
File diff suppressed because it is too large Load Diff
+4
View File
@@ -1,3 +1,7 @@
allowBuilds:
core-js: true
es5-ext: true
esbuild: true
onlyBuiltDependencies:
- esbuild
- '@tailwindcss/oxide'
+129
View File
@@ -0,0 +1,129 @@
#!/usr/bin/env bash
#
# Vendors foliate-js into src/lib/vendor/foliate-js/.
#
# foliate-js has no npm release and no build step; upstream recommends a git
# submodule. We copy instead, because this repo has no submodules (pdf.js is
# vendored the same way under static/pdfjs) and because a submodule would drag
# in 231 files / 13 MB, of which 191 files / 12 MB is a bundled pdf.js build we
# deliberately do not use — Chitai serves PDFs through static/pdfjs/web/viewer.html.
#
# Only the files reachable from view.js are copied: 15 upstream files, ~656 KB.
# pdf.js is NOT copied; a stub is written in its place (see below).
#
# To update: bump FOLIATE_SHA, re-run, review the diff, then smoke-test the
# reader — paginator.js is ~3800 lines of gesture and animation code and this
# fork is pushed to frequently.
#
# Usage: ./scripts/vendor-foliate.sh
set -euo pipefail
FOLIATE_REPO="https://github.com/readest/foliate-js.git"
FOLIATE_SHA="63a2eb1fc1e4813c4e849ccdb3d4be2c54a35869"
DEST="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)/src/lib/vendor/foliate-js"
TMP="$(mktemp -d)"
trap 'rm -rf "$TMP"' EXIT
# The reachable closure from view.js. Everything else upstream ships is either
# unreachable (dict.js, opds.js, footnotes.js, quote-image.js, uri-template.js),
# a demo (reader.js), or build tooling (rollup.config.js, eslint.config.js).
FILES=(
view.js
epub.js
epubcfi.js
paginator.js
fixed-layout.js
overlayer.js
progress.js
search.js
text-walker.js
tts.js
mobi.js
comic-book.js
fb2.js
vendor/zip.js
vendor/fflate.js
)
echo "Cloning $FOLIATE_REPO @ ${FOLIATE_SHA:0:7} ..."
git clone --quiet --filter=blob:none --no-checkout "$FOLIATE_REPO" "$TMP/foliate"
git -C "$TMP/foliate" checkout --quiet "$FOLIATE_SHA"
rm -rf "$DEST"
mkdir -p "$DEST/vendor"
for f in "${FILES[@]}"; do
if [ ! -f "$TMP/foliate/$f" ]; then
echo "ERROR: $f is missing upstream at ${FOLIATE_SHA:0:7}." >&2
echo "The file list in this script is stale; re-check the import graph." >&2
exit 1
fi
cp "$TMP/foliate/$f" "$DEST/$f"
done
cp "$TMP/foliate/LICENSE" "$DEST/LICENSE"
# view.js does `await import('./pdf.js')` inside makeBook. That is a static-string
# dynamic import, so Rollup resolves it at build time whether or not the code path
# ever runs — and upstream's pdf.js opens with `import '@pdfjs/pdf.min.mjs'`, a bare
# specifier that does not resolve here. Shipping this stub at that path keeps the
# build working without a Vite alias, and without vendoring 12 MB of pdf.js.
cat > "$DEST/pdf.js" <<'STUB'
// NOT upstream foliate-js. See README.chitai.md.
//
// Chitai renders PDFs with the pdf.js viewer vendored at static/pdfjs/, so
// foliate's PDF backend is not vendored. view.js still references this module
// from makeBook via a static-string dynamic import, which Rollup resolves at
// build time regardless of whether it executes — so the file has to exist.
//
// Throwing at module scope surfaces a legible message in the reader's error
// card if a PDF is ever routed to the EPUB reader by mistake, rather than a
// TypeError from `globalThis.pdfjsLib` being undefined.
throw new Error('foliate-js PDF rendering is not enabled in Chitai');
STUB
cat > "$DEST/README.chitai.md" <<EOF
# Vendored foliate-js
Do not edit these files. They are copied verbatim from upstream by
\`frontend/scripts/vendor-foliate.sh\`; local changes are lost on the next run.
| | |
| --- | --- |
| Upstream | <https://github.com/readest/foliate-js> (Readest's fork of johnfactotum/foliate-js) |
| Pinned commit | \`$FOLIATE_SHA\` |
| Licence | MIT — see \`LICENSE\` |
Readest's fork is used rather than upstream for its paginator work: touch/swipe
turn handling, fixed-layout spread centring, and a malformed-XHTML fallback in
\`loadDocument\`.
## What is here
Only the import closure reachable from \`view.js\`. Not vendored, because nothing
reaches them: \`dict.js\`, \`opds.js\`, \`footnotes.js\`, \`quote-image.js\`,
\`uri-template.js\`, \`reader.js\` (upstream's demo), and the build configs.
## pdf.js is ours, not upstream's
\`pdf.js\` in this directory is a **stub that throws**. Upstream's version imports
\`@pdfjs/pdf.min.mjs\` — a bare specifier backed by a 12 MB vendored pdf.js build —
and \`view.js\` reaches it through \`await import('./pdf.js')\`, which Rollup resolves
at build time even though Chitai never takes that path. Chitai serves PDFs from
\`static/pdfjs/web/viewer.html\` instead.
To enable foliate's PDF backend, add \`pdf.js\` and \`vendor/pdfjs/\` to the file list
in the vendor script and drop the stub.
## Updating
Bump \`FOLIATE_SHA\` in \`frontend/scripts/vendor-foliate.sh\`, re-run it, review the
diff, and smoke-test the reader — \`paginator.js\` is ~3800 lines of gesture and
animation code and this fork is pushed to frequently.
EOF
echo
echo "Vendored ${#FILES[@]} files + LICENSE + pdf.js stub + README.chitai.md to:"
echo " $DEST"
du -sh "$DEST" | sed 's/^/ /'
+2 -1
View File
@@ -4,7 +4,8 @@ pkgs.mkShell {
buildInputs = with pkgs; [
# Node.js ecosystem
nodejs_24
nodePackages.pnpm
pnpm
claude-code
];
shellHook = ''
+92 -62
View File
@@ -8,74 +8,100 @@
:root {
--radius: 0.625rem;
--background: oklch(1 0 0);
--foreground: oklch(0.129 0.042 264.695);
--card: oklch(1 0 0);
--card-foreground: oklch(0.129 0.042 264.695);
--popover: oklch(1 0 0);
--popover-foreground: oklch(0.129 0.042 264.695);
--primary: oklch(0.208 0.042 265.755);
--primary-foreground: oklch(0.984 0.003 247.858);
--secondary: oklch(0.968 0.007 247.896);
--secondary-foreground: oklch(0.208 0.042 265.755);
--muted: oklch(0.968 0.007 247.896);
--muted-foreground: oklch(0.554 0.046 257.417);
--accent: oklch(0.968 0.007 247.896);
--accent-foreground: oklch(0.208 0.042 265.755);
--destructive: oklch(0.577 0.245 27.325);
--border: oklch(0.929 0.013 255.508);
--input: oklch(0.929 0.013 255.508);
--ring: oklch(0.704 0.04 256.788);
--chart-1: oklch(0.646 0.222 41.116);
--chart-2: oklch(0.6 0.118 184.704);
--chart-3: oklch(0.398 0.07 227.392);
--chart-4: oklch(0.828 0.189 84.429);
--chart-5: oklch(0.769 0.188 70.08);
--sidebar: oklch(0.984 0.003 247.858);
--sidebar-foreground: oklch(0.129 0.042 264.695);
--sidebar-primary: oklch(0.208 0.042 265.755);
--sidebar-primary-foreground: oklch(0.984 0.003 247.858);
--sidebar-accent: oklch(0.968 0.007 247.896);
--sidebar-accent-foreground: oklch(0.208 0.042 265.755);
--sidebar-border: oklch(0.929 0.013 255.508);
--sidebar-ring: oklch(0.704 0.04 256.788);
/* Typography — system stacks, so nothing depends on a CDN or a webfont build. */
--app-font-sans: system-ui, -apple-system, 'Segoe UI', Roboto, 'Helvetica Neue', Arial, sans-serif;
--app-font-serif: Georgia, 'Iowan Old Style', 'Times New Roman', serif;
--app-font-mono: ui-monospace, 'SF Mono', 'Cascadia Mono', Menlo, Consolas, monospace;
/* Reading Room — light */
--background: #e9ebee;
--foreground: #1b1f24;
--card: #fbfbfc;
--card-foreground: #1b1f24;
--popover: #fbfbfc;
--popover-foreground: #1b1f24;
--primary: #1f5f5b;
--primary-foreground: #f2f7f6;
--secondary: #dfe3e8;
--secondary-foreground: #1b1f24;
--muted: #e2e5e9;
--muted-foreground: #7d858f;
--accent: #d7e5e3;
--accent-foreground: #164743;
--destructive: #a6402f;
--border: #d2d6dc;
--input: #d2d6dc;
--ring: #1f5f5b;
/* Semantic — deliberately not the accent, so state never reads as branding. */
--success: #2f7d4f;
--success-foreground: #f2f7f6;
--flag: #b08a1e;
--star: #c79a25;
--chart-1: #1f5f5b;
--chart-2: #2f7d4f;
--chart-3: #b08a1e;
--chart-4: #4c6b8a;
--chart-5: #a6402f;
--sidebar: #e2e5e9;
--sidebar-foreground: #1b1f24;
--sidebar-primary: #1f5f5b;
--sidebar-primary-foreground: #f2f7f6;
--sidebar-accent: #d7e5e3;
--sidebar-accent-foreground: #164743;
--sidebar-border: #d2d6dc;
--sidebar-ring: #1f5f5b;
}
.dark {
--background: oklch(0.129 0.042 264.695);
--foreground: oklch(0.984 0.003 247.858);
--card: oklch(0.208 0.042 265.755);
--card-foreground: oklch(0.984 0.003 247.858);
--popover: oklch(0.208 0.042 265.755);
--popover-foreground: oklch(0.984 0.003 247.858);
--primary: oklch(0.929 0.013 255.508);
--primary-foreground: oklch(0.208 0.042 265.755);
--secondary: oklch(0.279 0.041 260.031);
--secondary-foreground: oklch(0.984 0.003 247.858);
--muted: oklch(0.279 0.041 260.031);
--muted-foreground: oklch(0.704 0.04 256.788);
--accent: oklch(0.279 0.041 260.031);
--accent-foreground: oklch(0.984 0.003 247.858);
--destructive: oklch(0.704 0.191 22.216);
--border: oklch(1 0 0 / 10%);
--input: oklch(1 0 0 / 15%);
--ring: oklch(0.551 0.027 264.364);
--chart-1: oklch(0.488 0.243 264.376);
--chart-2: oklch(0.696 0.17 162.48);
--chart-3: oklch(0.769 0.188 70.08);
--chart-4: oklch(0.627 0.265 303.9);
--chart-5: oklch(0.645 0.246 16.439);
--sidebar: oklch(0.208 0.042 265.755);
--sidebar-foreground: oklch(0.984 0.003 247.858);
--sidebar-primary: oklch(0.488 0.243 264.376);
--sidebar-primary-foreground: oklch(0.984 0.003 247.858);
--sidebar-accent: oklch(0.279 0.041 260.031);
--sidebar-accent-foreground: oklch(0.984 0.003 247.858);
--sidebar-border: oklch(1 0 0 / 10%);
--sidebar-ring: oklch(0.551 0.027 264.364);
/* Reading Room — dark */
--background: #15191b;
--foreground: #e6eae9;
--card: #1d2226;
--card-foreground: #e6eae9;
--popover: #1d2226;
--popover-foreground: #e6eae9;
--primary: #6fbab0;
--primary-foreground: #0e1a19;
--secondary: #232a2d;
--secondary-foreground: #e6eae9;
--muted: #232a2d;
--muted-foreground: #78868a;
--accent: #1b3330;
--accent-foreground: #9fd8d0;
--destructive: #e0796b;
--border: #2a3034;
--input: #2a3034;
--ring: #6fbab0;
--success: #4fa97a;
--success-foreground: #0e1a19;
--flag: #d4a63a;
--star: #e5b84b;
--chart-1: #6fbab0;
--chart-2: #4fa97a;
--chart-3: #d4a63a;
--chart-4: #7f9dc0;
--chart-5: #e0796b;
--sidebar: #111517;
--sidebar-foreground: #e6eae9;
--sidebar-primary: #6fbab0;
--sidebar-primary-foreground: #0e1a19;
--sidebar-accent: #1b3330;
--sidebar-accent-foreground: #9fd8d0;
--sidebar-border: #2a3034;
--sidebar-ring: #6fbab0;
}
@theme inline {
--font-sans: var(--app-font-sans);
--font-serif: var(--app-font-serif);
--font-mono: var(--app-font-mono);
--radius-sm: calc(var(--radius) - 4px);
--radius-md: calc(var(--radius) - 2px);
--radius-lg: var(--radius);
@@ -95,6 +121,10 @@
--color-accent: var(--accent);
--color-accent-foreground: var(--accent-foreground);
--color-destructive: var(--destructive);
--color-success: var(--success);
--color-success-foreground: var(--success-foreground);
--color-flag: var(--flag);
--color-star: var(--star);
--color-border: var(--border);
--color-input: var(--input);
--color-ring: var(--ring);
+3 -1
View File
@@ -2,7 +2,8 @@
// for information about these interfaces
import type { ApiClient } from '$lib/server/api';
import type { User } from 'lucide-svelte';
import type { User } from '$lib/server/auth';
import type { ThemeConfig } from '$lib/theme/presets';
declare global {
namespace App {
@@ -11,6 +12,7 @@ declare global {
authToken: string | null;
api: ApiClient;
user: User;
theme: ThemeConfig;
}
// interface PageData {}
// interface PageState {}
+1
View File
@@ -4,6 +4,7 @@
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
%sveltekit.head%
<!--theme-->
</head>
<body data-sveltekit-preload-data="hover" class="overflow-hidden">
<div style="display: contents">%sveltekit.body%</div>
+26 -1
View File
@@ -1,5 +1,6 @@
import { ApiClient } from '$lib/server/api';
import { validateToken } from '$lib/server/auth';
import { THEME_COOKIE, parseThemeCookie, themeToCss } from '$lib/theme/presets';
import { redirect, type Handle } from '@sveltejs/kit';
import { sequence } from '@sveltejs/kit/hooks';
@@ -34,4 +35,28 @@ const protectedRoutesHandle: Handle = async ({ event, resolve }) => {
return resolve(event);
};
export const handle = sequence(authHandle, protectedRoutesHandle);
/**
* Inline the stored theme into the document head.
*
* The palette has to be in the very first byte of HTML the browser paints,
* otherwise every page load flashes the default theme before hydration swaps
* it. The `<!--theme-->` placeholder in app.html is the injection point.
*/
const themeHandle: Handle = async ({ event, resolve }) => {
const config = parseThemeCookie(event.cookies.get(THEME_COOKIE));
event.locals.theme = config;
return resolve(event, {
// The placeholder comment is kept, not replaced. Svelte 5 uses HTML
// comments as hydration markers, so SvelteKit warns when a chunk comes
// back with fewer comments than it went in with — removing this one is
// enough to trip that check.
transformPageChunk: ({ html }) =>
html.replace(
'<!--theme-->',
`<!--theme--><style id="chitai-theme">${themeToCss(config)}</style>`
)
});
};
export const handle = sequence(themeHandle, authHandle, protectedRoutesHandle);
+86 -18
View File
@@ -3,18 +3,36 @@ import { error } from '@sveltejs/kit';
import {
bookCoverUpload,
booksUpload,
bookIdsSchema,
bookQuerySchema,
deleteBookFilesSchema,
deleteBooksSchema,
editBookMetadataSchema,
updateBookProgressSchema,
duplicateDismissalSchema,
bookMergeSchema,
type Book,
type BooksUploadResult,
type DuplicateBookGroup,
bookFilesUpload
} from '$lib/schema/index';
import { stringCoerce, type PaginatedResponse } from '$lib/schema/common';
import { createQueryParams } from '$lib/utils';
/**
* The backend's own message for a failed response, rather than its JSON envelope.
*
* A refused duplicate answers 409 with a `detail` worth reading and the offending
* files in `extra`; passing the body through whole puts JSON in front of the reader.
*/
function detailOf(body: string): string {
try {
const parsed = JSON.parse(body);
return typeof parsed?.detail === 'string' ? parsed.detail : body;
} catch {
return body;
}
}
export const getBook = query(stringCoerce, async (id): Promise<Book> => {
const { locals } = getRequestEvent();
@@ -69,7 +87,9 @@ export const updateBookCover = form(bookCoverUpload, async ({ book_id, file }) =
return await response.json();
});
export const uploadBooks = form(booksUpload, async ({ library_id, files }) => {
export const uploadBooks = form(
booksUpload,
async ({ library_id, files }): Promise<BooksUploadResult> => {
const { locals } = getRequestEvent();
const formData = new FormData();
@@ -88,7 +108,8 @@ export const uploadBooks = form(booksUpload, async ({ library_id, files }) => {
}
return await response.json();
});
}
);
export const uploadBookFiles = form(bookFilesUpload, async ({ book_id, files }) => {
const { locals } = getRequestEvent();
@@ -101,8 +122,9 @@ export const uploadBookFiles = form(bookFilesUpload, async ({ book_id, files })
const response = await locals.api.postMultipart(`/books/${book_id}/files`, formData);
if (!response.ok) {
const message = await response.text();
error(response.status, message);
// 409 here means the file is already stored under a different book, which is
// something the reader can act on — so the message has to survive the trip.
error(response.status, detailOf(await response.text()));
}
return await response.json();
@@ -134,6 +156,65 @@ export const deleteBookFiles = command(deleteBookFilesSchema, async ({ book_id,
}
});
/**
* Books already in the library that look like copies of one another.
*
* Metadata only, so every group is a question rather than a verdict which is why
* the screen it feeds offers a way to disagree.
*/
export const listDuplicateBooks = query(
stringCoerce,
async (libraryId): Promise<DuplicateBookGroup[]> => {
const { locals } = getRequestEvent();
const response = await locals.api.get(`/books/duplicate-books?library_id=${libraryId}`);
if (!response.ok) error(response.status, detailOf(await response.text()));
return await response.json();
}
);
/**
* Fold several books into one and delete the records folded in.
*
* Irreversible, so the caller is expected to have shown what is about to happen.
*/
export const mergeBooks = command(
bookMergeSchema,
async ({ library_id, ...data }): Promise<Book> => {
const { locals } = getRequestEvent();
const response = await locals.api.post(`/books/merge?library_id=${library_id}`, data);
if (!response.ok) error(response.status, detailOf(await response.text()));
return await response.json();
}
);
/** Record that two books are not the same book, so the pair stops being proposed. */
export const dismissDuplicateBooks = command(duplicateDismissalSchema, async (data) => {
const { locals } = getRequestEvent();
const response = await locals.api.post('/books/duplicate-books/dismissals', data);
if (!response.ok) error(response.status, detailOf(await response.text()));
});
/** Undo a dismissal, so the pair is proposed again. */
export const restoreDuplicateBooks = command(duplicateDismissalSchema, async (data) => {
const { locals } = getRequestEvent();
const params = createQueryParams(data);
const response = await locals.api.delete(
`/books/duplicate-books/dismissals?${params.toString()}`
);
if (!response.ok) error(response.status, detailOf(await response.text()));
});
export const updateBookProgress = command(
updateBookProgressSchema,
async ({ book_ids, ...data }) => {
@@ -149,16 +230,3 @@ export const updateBookProgress = command(
}
}
);
export const markBooksAsComplete = command(bookIdsSchema, async (data) => {
const { locals } = getRequestEvent();
const params = createQueryParams(data);
const response = await locals.api.post(`/books/completed?${params.toString()}`, {});
if (!response.ok) {
const message = await response.text();
error(response.status, message);
}
});
@@ -0,0 +1,51 @@
import { command, getRequestEvent, query } from '$app/server';
import { error } from '@sveltejs/kit';
import { stringCoerce } from '$lib/schema/common';
import type { CalibreImport } from '$lib/schema/library';
/** The backend's own message for a failed response, rather than its JSON envelope. */
function detailOf(body: string): string {
try {
const parsed = JSON.parse(body);
return typeof parsed?.detail === 'string' ? parsed.detail : body;
} catch {
return body;
}
}
/**
* Where an import has got to.
*
* A `query` rather than a `command` so it can be refreshed, but it is polled on a timer
* rather than cached the answer changes on its own.
*
* Starting an import is deliberately **not** here: the archive goes straight to the
* backend through the proxy, so it never passes through this process. See the import
* screen's `upload`.
*/
export const getCalibreImport = query(stringCoerce, async (jobId): Promise<CalibreImport> => {
const { locals } = getRequestEvent();
const response = await locals.api.get(`/libraries/imports/${jobId}`);
if (!response.ok) error(response.status, detailOf(await response.text()));
return await response.json();
});
/**
* Ask an import to stop after the book it is on.
*
* Not an abort: a book abandoned mid-copy would leave files on disk with no row
* describing them. Whatever it has imported stays imported.
*/
export const cancelCalibreImport = command(stringCoerce, async (jobId): Promise<CalibreImport> => {
const { locals } = getRequestEvent();
const response = await locals.api.delete(`/libraries/imports/${jobId}`);
if (!response.ok) error(response.status, detailOf(await response.text()));
return await response.json();
});
+43
View File
@@ -0,0 +1,43 @@
import { command, form, getRequestEvent, query } from '$app/server';
import { createDeviceSchema, type Device } from '$lib/schema/device';
import { error } from '@sveltejs/kit';
import z from 'zod';
export const listDevices = query(async (): Promise<Device[]> => {
const { locals } = getRequestEvent();
const response = await locals.api.get(`/devices`);
if (!response.ok) error(500, 'An unkown error occurred');
const deviceResult = await response.json();
return deviceResult.items
});
export const createDevice = form(createDeviceSchema, async (data): Promise<Device> => {
const { locals } = getRequestEvent();
const response = await locals.api.post(`/devices`, data);
if (!response.ok) error(500, 'An unknown error occurred');
return await response.json();
});
export const regenerateDeviceApiKey = command(z.string(), async (deviceId): Promise<Device> => {
const { locals } = getRequestEvent();
const response = await locals.api.get(`/devices/${deviceId}/regenerate`);
if (!response.ok) error(500, 'An unknown error occurred');
return await response.json();
})
export const deleteDevice = command(z.string(), async (deviceId): Promise<void> => {
const { locals } = getRequestEvent();
const response = await locals.api.delete(`/devices/${deviceId}`);
if (!response.ok) error(500, 'An unknown error occurred');
});
+1
View File
@@ -2,6 +2,7 @@ export * from './auth.remote';
export * from './author.remote';
export * from './book.remote';
export * from './bookshelf.remote';
export * from './calibre-import.remote';
export * from './library.remote';
export * from './publisher.remote';
export * from './tag.remote';
@@ -0,0 +1,223 @@
<script lang="ts">
import { FileUp, Folder, FileText } from '@lucide/svelte';
import { Button } from '$lib/components/ui/button/index.js';
import type { FileRejectedReason } from '$lib/components/ui/file-drop-zone';
import { cn } from '$lib/utils';
let {
onUpload,
onFileRejected,
accept,
maxFileSize,
disabled = false,
class: className
}: {
onUpload: (files: File[]) => Promise<void> | void;
onFileRejected?: (opts: { reason: FileRejectedReason; file: File }) => void;
/** Comma separated extensions and/or MIME types, as the `accept` attribute takes. */
accept?: string;
/** Bytes. */
maxFileSize?: number;
disabled?: boolean;
class?: string;
} = $props();
/**
* Two inputs rather than one.
*
* `webkitdirectory` is not a filter — it switches the picker into folder mode,
* so a single input can offer files or folders but never both. The drop target
* has no such constraint and stays one area.
*/
let fileInput = $state<HTMLInputElement>();
let folderInput = $state<HTMLInputElement>();
let dragging = $state(false);
let busy = $state(false);
const active = $derived(!disabled && !busy);
function accepts(file: File): FileRejectedReason | undefined {
if (maxFileSize !== undefined && file.size > maxFileSize) return 'Maximum file size exceeded';
if (!accept) return undefined;
const name = file.name.toLowerCase();
const type = file.type.toLowerCase();
const ok = accept
.split(',')
.map((pattern) => pattern.trim().toLowerCase())
.some((pattern) => {
// Match on the pattern, not the file's type. Testing `type` here is
// what makes MOBI fail: browsers report no MIME type for it, so a
// ".mobi" rule never gets compared against the filename.
if (pattern.startsWith('.')) return name.endsWith(pattern);
if (pattern.endsWith('/*')) return type.startsWith(pattern.slice(0, -1));
return type === pattern;
});
return ok ? undefined : 'File type not allowed';
}
/** readEntries hands back at most 100 at a time and signals the end with an empty batch. */
function readAll(reader: FileSystemDirectoryReader): Promise<FileSystemEntry[]> {
return new Promise((resolve, reject) => {
const entries: FileSystemEntry[] = [];
const next = () =>
reader.readEntries((batch) => {
if (batch.length === 0) return resolve(entries);
entries.push(...batch);
next();
}, reject);
next();
});
}
/**
* Flattens a dropped entry into files, naming each one with its path inside the
* dropped folder so it matches what the folder picker puts in
* `webkitRelativePath` — which is what the upload form reads to keep structure.
*/
async function walk(entry: FileSystemEntry, prefix = ''): Promise<File[]> {
if (entry.isFile) {
const file = await new Promise<File>((resolve, reject) =>
(entry as FileSystemFileEntry).file(resolve, reject)
);
return [new File([file], `${prefix}${file.name}`, { type: file.type })];
}
if (entry.isDirectory) {
const entries = await readAll((entry as FileSystemDirectoryEntry).createReader());
const nested = await Promise.all(entries.map((e) => walk(e, `${prefix}${entry.name}/`)));
return nested.flat();
}
return [];
}
async function handleDrop(event: DragEvent) {
event.preventDefault();
dragging = false;
if (!active) return;
// Read the entries synchronously: DataTransfer is emptied as soon as this
// handler yields, so awaiting first loses everything that was dropped.
const entries = Array.from(event.dataTransfer?.items ?? [])
.filter((item) => item.kind === 'file')
.map((item) => item.webkitGetAsEntry())
.filter((entry): entry is FileSystemEntry => entry !== null);
// Older engines expose no entries; fall back to the flat list, which cannot
// contain folders anyway.
if (entries.length === 0) {
await submit(Array.from(event.dataTransfer?.files ?? []));
return;
}
busy = true;
try {
const nested = await Promise.all(entries.map((entry) => walk(entry)));
await submit(nested.flat());
} finally {
busy = false;
}
}
async function handleChange(event: Event) {
const input = event.currentTarget as HTMLInputElement;
const chosen = Array.from(input.files ?? []);
// Reset first so picking the same file twice still fires a change event.
input.value = '';
await submit(chosen);
}
async function submit(candidates: File[]) {
const accepted: File[] = [];
for (const file of candidates) {
const reason = accepts(file);
if (reason) onFileRejected?.({ file, reason });
else accepted.push(file);
}
if (accepted.length > 0) await onUpload(accepted);
}
</script>
<div
role="group"
aria-label="Add books"
aria-disabled={!active}
ondragover={(e) => {
e.preventDefault();
if (active) dragging = true;
}}
ondragleave={() => (dragging = false)}
ondrop={handleDrop}
class={cn(
'flex flex-col items-center gap-3 rounded-lg border-2 border-dashed border-border bg-accent/20 p-6 text-center transition-colors',
dragging && 'border-primary bg-accent/50',
!active && 'opacity-50',
className
)}
>
<div
class="flex size-12 place-items-center justify-center rounded-full border border-dashed border-border text-muted-foreground"
>
<FileUp class="size-5" />
</div>
<div class="flex flex-col gap-0.5">
<span class="font-medium">
{busy ? 'Reading folder…' : 'Drop books here'}
</span>
<span class="text-sm text-muted-foreground">A folder keeps its structure</span>
</div>
<span class="font-mono text-[10px] tracking-widest text-muted-foreground uppercase">or</span>
<div class="flex flex-wrap justify-center gap-2">
<Button
type="button"
variant="outline"
size="sm"
disabled={!active}
onclick={() => fileInput?.click()}
>
<FileText class="size-4" />
Choose files
</Button>
<Button
type="button"
variant="outline"
size="sm"
disabled={!active}
onclick={() => folderInput?.click()}
>
<Folder class="size-4" />
Choose folder
</Button>
</div>
<input
bind:this={fileInput}
type="file"
multiple
{accept}
class="hidden"
onchange={handleChange}
/>
<!-- webkitdirectory is why this needs to be a second input: it turns the
picker into a folder chooser rather than filtering what it accepts. -->
<input
bind:this={folderInput}
type="file"
multiple
webkitdirectory
class="hidden"
onchange={handleChange}
/>
</div>
@@ -1,161 +1,231 @@
<script lang="ts">
import { untrack } from 'svelte';
import { goto } from '$app/navigation';
import { resolve } from '$app/paths';
import { X } from '@lucide/svelte';
import * as Dialog from '$lib/components/ui/dialog/index.js';
import * as Field from '$lib/components/ui/field/index.js';
import { Button } from '$lib/components/ui/button';
import * as NativeSelect from '$lib/components/ui/native-select/index.js';
import {
displaySize,
FileDropZone,
type FileDropZoneProps
} from '$lib/components/ui/file-drop-zone';
import { X } from '@lucide/svelte';
import { toast } from 'svelte-sonner';
import { Button } from '$lib/components/ui/button';
import { Switch } from '$lib/components/ui/switch/index';
import { displaySize, type FileRejectedReason } from '$lib/components/ui/file-drop-zone';
import BookDropZone from './book-drop-zone.svelte';
import { tick } from 'svelte';
import { Spinner } from '$lib/components/ui/spinner/index';
import { getLibraryState } from '$lib/state/library.svelte';
import { uploadBooks } from '$lib/api';
import type { Book, PaginatedResponse } from '$lib/schema';
import { goto } from '$app/navigation';
import { getUploadQueueState } from '$lib/state/upload-queue.svelte';
let { open = $bindable() }: { open?: boolean } = $props();
let libraryState = getLibraryState();
$effect(() => {
uploadBooks.fields.library_id.set(libraryState.activeLibrary!.id);
});
let files = $derived(uploadBooks.fields.files.value() ?? []);
const libraryState = getLibraryState();
const queue = getUploadQueueState();
let libraryId = $state<number>(untrack(() => libraryState.activeLibrary!.id));
let files = $state<File[]>([]);
let autoUploadOnDrop = $state(true);
let navigateOnUpload = $state(true);
let formEl = $state<HTMLFormElement>();
const onUpload: FileDropZoneProps['onUpload'] = async (uploadedFiles) => {
uploadBooks.fields.files.set([...Array.from(files), ...uploadedFiles]);
if (autoUploadOnDrop && files.length > 0) {
await tick();
formEl?.requestSubmit();
const totalSize = $derived(files.reduce((sum, file) => sum + file.size, 0));
/**
* Collected rather than raised one at a time. A folder of a few hundred books
* carries covers and notes alongside them, and a toast per rejected file
* buries the screen.
*/
let rejected = $state<{ name: string; reason: FileRejectedReason }[]>([]);
let showRejected = $state(false);
// Start each visit clean — a list left over from last time reads as though it
// applies to what was just chosen.
$effect(() => {
if (open) {
untrack(() => {
files = [];
rejected = [];
showRejected = false;
libraryId = libraryState.activeLibrary!.id;
});
}
});
const onUpload = async (uploadedFiles: File[]) => {
// Rename to the path relative to the chosen folder, which is what decides
// how the API groups files into books. Dropped folders already arrive named
// this way, since the drop zone builds the path while walking them.
const named = uploadedFiles.map(
(file) => new File([file], file.webkitRelativePath || file.name, { type: file.type })
);
// Same relative path twice is the same file — dropping a folder a second
// time should not queue everything again.
const seen = new Set(files.map((file) => file.name));
files = [...files, ...named.filter((file) => !seen.has(file.name))];
if (autoUploadOnDrop) await startUpload();
};
const onFileRejected: FileDropZoneProps['onFileRejected'] = async ({ reason, file }) => {
toast.error(`${file.name} failed to upload!`, { description: reason });
const onFileRejected = ({ reason, file }: { reason: FileRejectedReason; file: File }) => {
rejected = [...rejected, { name: file.webkitRelativePath || file.name, reason }];
};
function navigateToBooks(books: PaginatedResponse<Book>) {
/**
* Files picked from a folder carry their whole relative path as the name, so
* truncating the end would cut off the filename — the only part worth reading.
* Split it and let the folder sit on its own, quieter line.
*/
function splitPath(path: string) {
const cut = path.lastIndexOf('/');
return cut === -1
? { dir: '', name: path }
: { dir: path.slice(0, cut), name: path.slice(cut + 1) };
}
/**
* Hands the books to the queue and closes.
*
* Nothing is awaited here: the queue lives in the root layout and reports
* through the tray, so the import carries on while the library stays usable.
*/
function startUpload() {
if (files.length === 0) return;
const queued = files;
const target = libraryId;
files = [];
rejected = [];
open = false;
let libraryId = books.items[0].library_id;
libraryState.setActive(libraryId);
if (books.items.length === 1) {
goto(`/book/${books.items[0].id}`);
} else {
goto(`/library/${libraryId}/view?orderBy=created_at&sortOrder=desc`);
queue.enqueue(target, queued, ({ created, firstBook }) => {
if (created === 0) return;
const library = libraryState.libraries.find((lib) => lib.id === target);
if (library) library.total = (library.total ?? 0) + created;
// Only for a single book. Jumping somewhere after a bulk import would
// land minutes after the reader moved on.
if (navigateOnUpload && created === 1 && firstBook) {
libraryState.setActive(target);
void goto(resolve('/(root)/(library)/book/[bookId]', { bookId: String(firstBook.id) }));
}
});
}
</script>
<Dialog.Root bind:open>
<Dialog.Content>
{#if uploadBooks.pending}
<div class="flex flex-col items-center gap-4">
<span class="text-lg font-semibold"
>Uploading {uploadBooks.fields.files.value().length} files...</span
>
<Spinner class="scale-150" />
</div>
{:else}
<!--
Wider than the default lg: a folder's worth of rows needs the room.
overflow-hidden and the min-w-0 on the body below are what keep a long
filename inside the dialog. Dialog.Content is a grid, and grid and flex
items default to min-width:auto — they refuse to shrink below their
content's intrinsic width, so one long name widened the body and pushed it
straight through the dialog's edge regardless of any truncate further down.
-->
<Dialog.Content class="overflow-hidden sm:max-w-2xl">
<Dialog.Header>
<Dialog.Title>Upload Books</Dialog.Title>
<Dialog.Title>Add books</Dialog.Title>
<Dialog.Description>
Files or a folder. A folder becomes one book per directory.
</Dialog.Description>
</Dialog.Header>
<form
{...uploadBooks.enhance(async ({ submit, form }) => {
try {
await submit();
// Check if there are any validation issues
const issues = uploadBooks.fields.allIssues();
if (issues && issues.length > 0) {
return;
}
// Update library book count
const count = uploadBooks.result.total
libraryState.libraries.find(lib => uploadBooks.fields.library_id.value() == lib.id.toString())!.total += count
// Reset the files field
uploadBooks.fields.files.set([]);
toast.success('Books successfully uploaded!');
if (navigateOnUpload) {
navigateToBooks(uploadBooks.result);
}
} catch (error) {
console.error('Failed to upload book: ', error);
toast.error('Failed to upload books');
}
})}
bind:this={formEl}
enctype="multipart/form-data"
class="flex w-full flex-col gap-2 p-4"
>
<!-- Library select field -->
<Field.Label for="library_id">Select Library</Field.Label>
<NativeSelect.Root {...uploadBooks.fields.library_id.as('select')} class="w-36">
{#each libraryState.libraries as library}
<NativeSelect.Option value={library.id}>
{library.name}
</NativeSelect.Option>
<div class="flex w-full min-w-0 flex-col gap-3">
<div class="flex flex-col gap-1.5">
<Field.Label for="library_id">Library</Field.Label>
<NativeSelect.Root id="library_id" bind:value={libraryId} class="w-48">
{#each libraryState.libraries as library (library.id)}
<NativeSelect.Option value={library.id}>{library.name}</NativeSelect.Option>
{/each}
</NativeSelect.Root>
</div>
<FileDropZone
<BookDropZone
{onUpload}
{onFileRejected}
directory={true}
accept=".pdf,.epub,.mobi,application/pdf,application/epub+zip,application/x-mobipocket-ebook"
sublabel="Only PDF, EPUB, and MOBI files supported"
/>
<input class="hidden" {...uploadBooks.fields.files.as('file multiple')} />
<div class="flex max-h-[300px] flex-col gap-2 overflow-y-auto">
{#each files as file, idx}
<div class="flex place-items-center justify-between gap-2">
<div class="flex flex-col">
<span>{file.name}</span>
<span class="text-xs text-muted-foreground">{displaySize(file.size)}</span>
{#if files.length > 0}
<div class="flex items-baseline justify-between border-b pb-1 text-sm">
<span><strong class="tabular-nums">{files.length}</strong> ready to upload</span>
<span class="font-mono text-xs text-muted-foreground tabular-nums">
{displaySize(totalSize)}
</span>
</div>
{/if}
<div class="flex max-h-[300px] min-w-0 flex-col gap-2 overflow-y-auto">
{#each files as file, idx (file.name)}
{@const location = splitPath(file.name)}
<div class="flex min-w-0 items-center justify-between gap-2">
<!-- flex-1 as well as min-w-0: without a constrained width there is
nothing for truncate to act against and the row pushes the dialog wide -->
<div class="flex min-w-0 flex-1 flex-col">
<span class="truncate text-sm" title={file.name}>{location.name}</span>
<span class="flex min-w-0 items-baseline gap-2 text-xs text-muted-foreground">
{#if location.dir}
<span class="truncate font-mono" title={location.dir}>{location.dir}</span>
{/if}
<span class="shrink-0 tabular-nums">{displaySize(file.size)}</span>
</span>
</div>
<Button
variant="outline"
size="icon"
onclick={() => {
uploadBooks.fields.files.set([
...Array.from(files).slice(0, idx),
...Array.from(files).slice(idx + 1)
]);
}}
class="shrink-0"
onclick={() => (files = files.filter((_, i) => i !== idx))}
>
<X />
<span class="sr-only">Remove {file.name}</span>
</Button>
</div>
{/each}
</div>
<div class="flex flex-col gap-2">
<div class="flex items-center gap-2">
<Switch bind:checked={autoUploadOnDrop} />
<Field.Label for="auto-upload-on-drop">Auto upload on file drop</Field.Label>
<Button type="submit" class="ml-auto w-fit">Upload</Button>
{#if rejected.length > 0}
<div class="rounded-md border border-star/50 bg-star/10 p-2 text-sm">
<div class="flex items-center justify-between gap-2">
<span>
{rejected.length}
{rejected.length === 1 ? 'file was' : 'files were'} skipped
</span>
<Button variant="ghost" size="sm" onclick={() => (showRejected = !showRejected)}>
{showRejected ? 'Hide' : 'Show'}
</Button>
</div>
<div class="flex items-center gap-2">
<Switch bind:checked={navigateOnUpload} />
<Field.Label for="navigate-to-book">Navigate to book on upload</Field.Label>
</div>
</div>
</form>
{#if showRejected}
<ul class="mt-2 flex max-h-32 min-w-0 flex-col gap-1 overflow-y-auto">
{#each rejected as entry (entry.name)}
<li class="flex min-w-0 justify-between gap-2 text-xs text-muted-foreground">
<span class="min-w-0 flex-1 truncate" title={entry.name}>
{splitPath(entry.name).name}
</span>
<span class="shrink-0">{entry.reason}</span>
</li>
{/each}
</ul>
{/if}
</div>
{/if}
<div class="flex flex-col gap-2 border-t pt-3">
<div class="flex items-center gap-2">
<Switch id="auto-upload-on-drop" bind:checked={autoUploadOnDrop} />
<Field.Label for="auto-upload-on-drop" class="font-normal">
Start as soon as books are added
</Field.Label>
<Button class="ml-auto w-fit" disabled={files.length === 0} onclick={startUpload}>
Upload
</Button>
</div>
<div class="flex items-center gap-2">
<Switch id="navigate-to-book" bind:checked={navigateOnUpload} />
<Field.Label for="navigate-to-book" class="font-normal">
Open the book when a single one is added
</Field.Label>
</div>
</div>
</div>
</Dialog.Content>
</Dialog.Root>
@@ -0,0 +1,50 @@
<script lang="ts">
import { createDevice } from '$lib/api/device.remote';
import * as Field from '$lib/components/ui/field/index.js';
import { Input } from '$lib/components/ui/input/index.js';
import { Button } from '$lib/components/ui/button/index.js';
import { toast } from 'svelte-sonner';
interface Props {
onSuccess?: () => void;
}
let { onSuccess }: Props = $props();
</script>
<form
{...createDevice.enhance(async ({ form, submit }) => {
try {
await submit();
const issues = createDevice.fields.allIssues();
if (issues && issues.length > 0) {
return;
}
const deviceName = createDevice.fields.name.value();
form.reset();
toast.success(`Device '${deviceName}' created`);
onSuccess?.();
} catch (error) {
console.error('Failed to create device: ', error);
toast.error('Failed to create device');
}
})}
class="flex flex-col gap-3"
>
<Field.Set>
<Field.Group>
<Field.Field>
<Field.Label for="name">Device name</Field.Label>
<Input {...createDevice.fields.name.as('text')} placeholder="e.g. Kindle Paperwhite" />
{#each createDevice.fields.name.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
</Field.Group>
</Field.Set>
<Button type="submit" class="ml-auto w-24">Create</Button>
</form>
@@ -1,39 +1,68 @@
<script lang="ts">
import * as Dialog from '$lib/components/ui/dialog/index.js';
import * as Tabs from '$lib/components/ui/tabs/index.js';
import { Button } from '$lib/components/ui/button/index.js';
import { type Book } from '$lib/schema';
import EditCover from './edit-cover.svelte';
import EditFiles from './edit-files.svelte';
import EditMetadata from './edit-metadata.svelte';
let { book, open = $bindable() }: { book?: Book; open: boolean } = $props();
/**
* The metadata form lives in a child, but Save belongs in the dialog footer —
* a submit button reaches it by id rather than the footer having to duplicate
* the submit logic.
*/
const METADATA_FORM_ID = 'edit-book-metadata';
</script>
<!--
This dialog is mounted once, in (root)/(library)/+layout.svelte, and `book`
changes underneath it as different books are edited. bookToEdit is never
cleared, so {#if book} stays true and the forms never unmount — they seed
local state from `book` on mount, so without this key you would open Edit on a
second book and see the first book's authors, tags and cover, then save them
onto the wrong record. Keying on the id remounts the forms per book.
-->
<Dialog.Root bind:open>
{#if book}
<Dialog.Content class="sm:max-w-xl">
<Tabs.Root value="metadata" class="h-[500px] max-w-xl py-4 md:h-[700px]">
<Tabs.List class="grid w-full grid-cols-3">
<Tabs.Trigger value="metadata">Metadata</Tabs.Trigger>
<Tabs.Trigger value="cover">Cover</Tabs.Trigger>
<Tabs.Trigger value="files">Files</Tabs.Trigger>
</Tabs.List>
{#key book.id}
<Dialog.Content
class="flex h-[85vh] max-w-[min(72rem,95vw)] flex-col gap-0 overflow-hidden p-0 sm:max-w-[min(72rem,95vw)]"
>
<Dialog.Header class="shrink-0 space-y-0 border-b px-5 py-3 text-left">
<Dialog.Title class="truncate font-serif text-base font-normal">{book.title}</Dialog.Title
>
<Dialog.Description class="truncate text-xs">
{book.authors.map((author) => author.name).join(', ') || 'Unknown author'}
</Dialog.Description>
</Dialog.Header>
<!-- Metadata form -->
<Tabs.Content value="metadata" class="h-full overflow-y-auto pb-1">
<EditMetadata {book} {open} />
</Tabs.Content>
<!-- Cover form -->
<Tabs.Content value="cover">
<EditCover {book} {open} />
</Tabs.Content>
<!-- Add files form -->
<Tabs.Content value="files">
<!-- Rail and form scroll independently so the footer never moves -->
<div class="grid min-h-0 flex-1 grid-cols-1 md:grid-cols-[16rem_1fr]">
<aside
class="flex min-w-0 flex-col gap-5 overflow-y-auto border-b bg-sidebar p-4 md:border-r md:border-b-0"
>
<EditCover {book} />
<EditFiles {book} />
</Tabs.Content>
</Tabs.Root>
</aside>
<div class="min-h-0 overflow-y-auto p-5">
<EditMetadata {book} bind:open formId={METADATA_FORM_ID} />
</div>
</div>
<Dialog.Footer class="shrink-0 items-center gap-2 border-t px-5 py-3 sm:justify-between">
<!-- Cover and file changes hit the server as they happen, while the
fields wait for Save. Saying so is the cheapest way to stop
Cancel reading as "undo everything". -->
<p class="text-xs text-muted-foreground">Cover and file changes apply immediately</p>
<div class="flex gap-2">
<Button variant="outline" onclick={() => (open = false)}>Cancel</Button>
<Button type="submit" form={METADATA_FORM_ID}>Save changes</Button>
</div>
</Dialog.Footer>
</Dialog.Content>
{/key}
{/if}
</Dialog.Root>
@@ -1,33 +1,23 @@
<script lang="ts">
import {
FileDropZone,
type FileDropZoneProps,
displaySize
} from '$lib/components/ui/file-drop-zone/index.js';
import { Button } from '$lib/components/ui/button/index.js';
import * as Field from '$lib/components/ui/field/index.js';
import { Switch } from '$lib/components/ui/switch/index';
import BookImage from '$lib/components/view/book-image.svelte';
import { tick, untrack } from 'svelte';
import { toast } from 'svelte-sonner';
import { updateBookCover } from '$lib/api';
import type { Book } from '$lib/schema';
import { tick, untrack } from 'svelte';
import { X } from '@lucide/svelte';
import BookImage from '$lib/components/view/book-image.svelte';
import { FileDropZone, type FileDropZoneProps } from '$lib/components/ui/file-drop-zone/index.js';
let { book, open = $bindable() }: { book: Book; open: boolean } = $props();
let { book }: { book: Book } = $props();
let formEl = $state<HTMLFormElement>();
let coverImagePreview = $state(`/api/${book.cover_image}`);
let autoUploadOnDrop = $state(true);
// Seeded once per mount — see the key in edit-book.svelte
let coverImagePreview = $state<string>(untrack(() => `/api/${book.cover_image}`));
const onUpload: FileDropZoneProps['onUpload'] = async (uploadedFiles) => {
updateBookCover.fields.file.set(uploadedFiles[0]);
updateCoverPreview();
if (autoUploadOnDrop && updateBookCover.fields.file.value()) {
await tick();
formEl?.requestSubmit();
}
};
function updateCoverPreview() {
@@ -35,14 +25,14 @@
if (file && file.type.startsWith('image/')) {
const reader = new FileReader();
reader.onloadend = () => {
coverImagePreview = reader.result;
coverImagePreview = reader.result as string;
};
reader.readAsDataURL(file);
}
}
const onFileRejected: FileDropZoneProps['onFileRejected'] = async ({ reason, file }) => {
toast.error(`${file.name} failed to upload!`, { description: reason });
toast.error(`${file.name} was not used`, { description: reason });
};
$effect(() => {
@@ -53,65 +43,40 @@
});
</script>
<section class="flex min-w-0 flex-col gap-2">
<h3 class="font-mono text-[10px] tracking-widest text-muted-foreground uppercase">Cover</h3>
<BookImage src={coverImagePreview} class="w-full rounded-md border" />
<form
bind:this={formEl}
{...updateBookCover.enhance(async ({ submit, form }) => {
try {
await submit();
form.reset();
open = false;
toast.success('Updated book cover!');
// Deliberately does not close the dialog. The cover is one panel of a
// larger form now, and closing here would throw away metadata edits
// the reader has not saved yet.
toast.success('Cover updated');
} catch (error) {
console.error('Failed to update book cover: ', error);
toast.error('Failed to update cover.');
toast.error('Failed to update the cover');
}
})}
enctype="multipart/form-data"
class="grid grid-cols-[1fr_2fr] gap-4 p-6"
class="flex flex-col gap-2"
>
<input class="hidden" {...updateBookCover.fields.book_id.as('text')} />
<input class="hidden" {...updateBookCover.fields.file.as('file')} />
<div class="flex flex-col gap-2">
<BookImage src={coverImagePreview} class="w-64 rounded" />
</div>
<div class="flex flex-col gap-2">
<FileDropZone
{onUpload}
{onFileRejected}
accept=".jpeg,.jpg,.png,.webp,image/*"
label="Only JPEG, PNG, and WEBP images supported"
label="Replace cover"
sublabel="JPEG, PNG or WEBP"
maxFiles={1}
fileCount={updateBookCover.fields.file.value() ? 1 : 0}
/>
<input class="hidden" {...updateBookCover.fields.file.as('file')} />
<div class="flex flex-col gap-2">
{#if updateBookCover.fields.file.value()}
<div class="flex place-items-center justify-between gap-2">
<div class="flex flex-col">
<span>{updateBookCover.fields.file.value().name}</span>
<span class="text-xs text-muted-foreground"
>{displaySize(updateBookCover.fields.file.value().size)}</span
>
</div>
<Button
variant="outline"
size="icon"
onclick={() => {
updateBookCover.fields.file.set(undefined);
coverImagePreview = `/api/${book.cover_image}`;
}}
>
<X />
</Button>
</div>
{/if}
</div>
<div class="flex flex-row items-center space-x-2">
<Switch bind:checked={autoUploadOnDrop} />
<Field.Label for="auto-upload-on-drop">Auto upload on file drop</Field.Label>
<Button type="submit" class="ml-auto w-fit" disabled={!updateBookCover.fields.file.value()}>Upload</Button>
</div>
</div>
</form>
</section>
@@ -1,98 +1,213 @@
<script lang="ts">
import { uploadBookFiles } from '$lib/api';
import {
displaySize,
FileDropZone,
type FileDropZoneProps
} from '$lib/components/ui/file-drop-zone';
import { Button } from '$lib/components/ui/button';
import * as Field from '$lib/components/ui/field/index.js';
import { Switch } from '$lib/components/ui/switch/index';
import { X } from '@lucide/svelte';
import type { Book } from '$lib/schema';
import { tick } from 'svelte';
import { tick, untrack } from 'svelte';
import { toast } from 'svelte-sonner';
import { Download, Trash2 } from '@lucide/svelte';
import { uploadBookFiles } from '$lib/api';
import type { Book, BookFile } from '$lib/schema';
import { formatFileSize, getFileType } from '$lib/utils';
import { getBookOperationsState } from '$lib/state/bookOperations.svelte';
import * as AlertDialog from '$lib/components/ui/alert-dialog/index.js';
import * as Field from '$lib/components/ui/field/index.js';
import { Button, buttonVariants } from '$lib/components/ui/button/index.js';
import { Checkbox } from '$lib/components/ui/checkbox/index.js';
import { FileDropZone, type FileDropZoneProps } from '$lib/components/ui/file-drop-zone';
import { Spinner } from '$lib/components/ui/spinner/index';
let { book }: { book: Book } = $props();
let files = $derived(uploadBookFiles.fields.files.value() ?? []);
const bookOps = getBookOperationsState();
let autoUploadOnDrop = $state(true);
/**
* The book's files, owned locally.
*
* `book` is a snapshot handed down from bookOperations rather than a live
* query, so invalidating after a delete does not reach it. Keeping the list
* here lets the rail reflect an add or a remove straight away. Seeded once per
* mount — edit-book.svelte keys the dialog on book.id.
*/
let files = $state<BookFile[]>(untrack(() => [...book.files]));
// The field's value is a sparse-ish list until the form settles, so narrow it
// before rendering rather than asserting at each use.
let pending = $derived(
(uploadBookFiles.fields.files.value() ?? []).filter((file): file is File => Boolean(file))
);
let fileToDelete = $state<BookFile>();
let deleteFromDisk = $state(true);
let confirmOpen = $state(false);
let formEl = $state<HTMLFormElement>();
const onUpload: FileDropZoneProps['onUpload'] = async (uploadedFiles) => {
uploadBookFiles.fields.files.set([...Array.from(files), ...uploadedFiles]);
if (autoUploadOnDrop && files.length > 0) {
const onUpload: FileDropZoneProps['onUpload'] = async (uploaded) => {
uploadBookFiles.fields.files.set([...Array.from(pending), ...uploaded]);
await tick();
formEl?.requestSubmit();
}
};
const onFileRejected: FileDropZoneProps['onFileRejected'] = async ({ reason, file }) => {
toast.error(`${file.name} failed to upload!`, { description: reason });
toast.error(`${file.name} was not added`, { description: reason });
};
/**
* The API's own words, when it has any.
*
* Adding a file the library already holds under another book is refused with a
* 409 naming it — far more use than "failed to add files". SvelteKit hands an
* `error()` back as an HttpError on the client, so the message sits on `body`.
*/
function apiMessage(error: unknown): string | undefined {
if (typeof error !== 'object' || error === null) return undefined;
const body = (error as { body?: { message?: string } }).body;
if (typeof body?.message === 'string') return body.message;
return error instanceof Error ? error.message : undefined;
}
function confirmDelete(file: BookFile) {
fileToDelete = file;
confirmOpen = true;
}
async function removeFile() {
if (!fileToDelete) return;
const target = fileToDelete;
confirmOpen = false;
// Drop it from the list first: the request invalidates the books query, but
// this dialog holds its own copy of the book and would not see that.
files = files.filter((file) => file.id !== target.id);
await bookOps.deleteBookFiles(book.id, [target.id], deleteFromDisk);
fileToDelete = undefined;
}
</script>
<section class="flex min-w-0 flex-col gap-2">
<h3 class="font-mono text-[10px] tracking-widest text-muted-foreground uppercase">Files</h3>
{#if files.length === 0}
<p class="text-xs text-muted-foreground">
No files yet. Add one below so this book can be read or downloaded.
</p>
{/if}
<ul class="flex flex-col gap-1.5">
{#each files as file (file.id)}
<li class="flex items-center gap-2 rounded-md border bg-background p-2">
<span
class="rounded-sm bg-accent px-1.5 py-0.5 font-mono text-[9px] font-semibold text-accent-foreground"
>
{getFileType(file.filename)}
</span>
<span class="min-w-0 flex-1">
<span class="block truncate text-xs" title={file.filename}>{file.filename}</span>
<span class="block font-mono text-[10px] text-muted-foreground tabular-nums">
{formatFileSize(file.size)}
</span>
</span>
<Button
variant="ghost"
size="icon"
class="size-7 shrink-0"
title="Download {file.filename}"
onclick={() => bookOps.downloadBookFile(book.id, file.id, file.filename)}
>
<Download class="size-3.5" />
<span class="sr-only">Download {file.filename}</span>
</Button>
<Button
variant="ghost"
size="icon"
class="size-7 shrink-0 text-destructive hover:text-destructive"
title="Remove {file.filename}"
onclick={() => confirmDelete(file)}
>
<Trash2 class="size-3.5" />
<span class="sr-only">Remove {file.filename}</span>
</Button>
</li>
{/each}
{#each pending as file (file.name)}
<li
class="flex items-center gap-2 rounded-md border border-dashed bg-background p-2 text-muted-foreground"
>
<Spinner class="size-3.5 shrink-0" />
<span class="min-w-0 flex-1">
<span class="block truncate text-xs">{file.name}</span>
<span class="block font-mono text-[10px] tabular-nums">Uploading…</span>
</span>
</li>
{/each}
</ul>
<form
{...uploadBookFiles.enhance(async ({ submit, form }) => {
bind:this={formEl}
{...uploadBookFiles.enhance(async ({ submit }) => {
try {
await submit();
// Check if there are any validation issues
const issues = uploadBookFiles.fields.allIssues();
if (issues && issues.length > 0) {
return;
}
if (issues && issues.length > 0) return;
// Reset the files field
// The endpoint answers with the updated book, so the new files come
// back with their ids rather than having to be guessed at.
files = uploadBookFiles.result?.files ?? files;
uploadBookFiles.fields.files.set([]);
toast.success('Files successfully added!');
toast.success('Files added');
} catch (error) {
console.error('Failed to upload files: ', error);
toast.error('Failed to upload files');
console.error('Failed to add files: ', error);
toast.error(apiMessage(error) ?? 'Failed to add files');
uploadBookFiles.fields.files.set([]);
}
})}
bind:this={formEl}
enctype="multipart/form-data"
class="flex w-full flex-col gap-2 p-4"
class="flex flex-col gap-2"
>
<input {...uploadBookFiles.fields.book_id.as('hidden', book.id)} />
<input class="hidden" {...uploadBookFiles.fields.files.as('file multiple')} />
<FileDropZone
{onUpload}
{onFileRejected}
accept=".pdf,.epub,.mobi,application/pdf,application/epub+zip,application/x-mobipocket-ebook"
sublabel="Only PDF, EPUB, and MOBI files supported"
label="Add a file"
sublabel="EPUB, PDF or MOBI"
/>
<input class="hidden" {...uploadBookFiles.fields.files.as('file multiple')} />
<div class="flex flex-col gap-2">
{#each files as file, idx}
<div class="flex place-items-center justify-between gap-2">
<div class="flex flex-col">
<span>{file.name}</span>
<span class="text-xs text-muted-foreground">{displaySize(file.size)}</span>
</div>
<Button
variant="outline"
size="icon"
onclick={() => {
uploadBookFiles.fields.files.set([
...Array.from(files).slice(0, idx),
...Array.from(files).slice(idx + 1)
]);
}}
>
<X />
</Button>
</div>
{/each}
</form>
</section>
<AlertDialog.Root bind:open={confirmOpen}>
<AlertDialog.Content>
<AlertDialog.Header>
<AlertDialog.Title>Remove {fileToDelete?.filename}?</AlertDialog.Title>
<AlertDialog.Description>
This cannot be undone. The other files on this book are not affected.
</AlertDialog.Description>
</AlertDialog.Header>
<!-- The API takes these as separate outcomes: drop the record, or drop the
record and the file on disk. Leaving it implicit would mean deleting
someone's only copy without saying so. -->
<div class="flex items-center gap-2">
<Checkbox id="delete-from-disk" bind:checked={deleteFromDisk} />
<Field.Label for="delete-from-disk" class="font-normal">
Also delete the file from the filesystem
</Field.Label>
</div>
<div class="flex flex-row items-center space-x-2">
<Switch bind:checked={autoUploadOnDrop} />
<Field.Label for="auto-upload-on-drop">Auto upload on file drop</Field.Label>
<Button type="submit" class="ml-auto w-fit">Upload</Button>
</div>
</form>
<AlertDialog.Footer>
<AlertDialog.Cancel>Cancel</AlertDialog.Cancel>
<AlertDialog.Action class={buttonVariants({ variant: 'destructive' })} onclick={removeFile}>
Remove
</AlertDialog.Action>
</AlertDialog.Footer>
</AlertDialog.Content>
</AlertDialog.Root>
@@ -1,26 +1,32 @@
<script lang="ts">
import * as Card from '$lib/components/ui/card/index.js';
import * as Field from '$lib/components/ui/field/index.js';
import { TagsInput, type TagsInputProps } from '$lib/components/ui/tags-input/index.js';
import { Input } from '$lib/components/ui/input/index.js';
import { Textarea } from '$lib/components/ui/textarea/index.js';
import { Button } from '$lib/components/ui/button/index.js';
import { untrack } from 'svelte';
import { toast } from 'svelte-sonner';
import { Minus, Plus } from '@lucide/svelte';
import { updateBookMetadata } from '$lib/api';
import type { Book } from '$lib/schema';
import { untrack } from 'svelte';
import { Minus, Plus } from '@lucide/svelte';
let { book, open = $bindable() }: { book: Book; open: boolean } = $props();
import * as Field from '$lib/components/ui/field/index.js';
import { Button } from '$lib/components/ui/button/index.js';
import { Input } from '$lib/components/ui/input/index.js';
import { TagsInput, type TagsInputProps } from '$lib/components/ui/tags-input/index.js';
import { Textarea } from '$lib/components/ui/textarea/index.js';
let authors = $state(book.authors.map((author) => author.name) || []);
let tags = $state(book.tags.map((tag) => tag.name) || []);
let identifierKeys = $state(Object.keys(book.identifiers));
let identifierValues = $state(Object.values(book.identifiers));
let {
book,
open = $bindable(),
/** Lets the dialog footer own the submit button via the `form` attribute. */
formId
}: { book: Book; open: boolean; formId: string } = $props();
// Seeded once per mount; edit-book.svelte keys this form on book.id so a
// different book gets a fresh form rather than the previous book's values.
let authors = $state(untrack(() => book.authors.map((author) => author.name) || []));
let tags = $state(untrack(() => book.tags.map((tag) => tag.name) || []));
let identifierKeys = $state(untrack(() => Object.keys(book.identifiers)));
let identifierValues = $state(untrack(() => Object.values(book.identifiers)));
function handleAddIdentifier() {
// Add empty strings to both arrays
identifierKeys = [...identifierKeys, ''];
identifierValues = [...identifierValues, ''];
}
@@ -86,9 +92,17 @@
});
</script>
<Card.Root class="w-full ">
{#snippet groupHeading(label: string)}
<h3
class="col-span-full mt-2 border-b pb-1 font-mono text-[10px] tracking-widest text-primary uppercase first:mt-0"
>
{label}
</h3>
{/snippet}
<form
{...updateBookMetadata.enhance(async ({ submit, form }) => {
id={formId}
{...updateBookMetadata.enhance(async ({ submit }) => {
try {
await submit();
@@ -99,28 +113,19 @@
}
open = false;
book = book;
toast.success('Updated book metadata!');
} catch (error) {
console.error('Error occurred updating book metadata: ', error);
toast.error('Failed to update book metadata.');
}
})}
class="grid grid-cols-1 items-start gap-x-4 gap-y-3 sm:grid-cols-2"
>
<Card.Content>
<Field.Set>
<Field.Group class="flex flex-col">
<!-- Book ID field -->
<Field.Field class="hidden">
<Field.Label for="book_id">Book ID</Field.Label>
<Input {...updateBookMetadata.fields.book_id.as('text')} />
{#each updateBookMetadata.fields.book_id.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
<input class="hidden" {...updateBookMetadata.fields.book_id.as('text')} />
<!-- Title field -->
<Field.Field>
{@render groupHeading('Identity')}
<Field.Field class="col-span-full">
<Field.Label for="title">Title</Field.Label>
<Input {...updateBookMetadata.fields.title.as('text')} />
{#each updateBookMetadata.fields.title.issues() ?? [] as issue}
@@ -128,7 +133,6 @@
{/each}
</Field.Field>
<!-- Subtitle field -->
<Field.Field>
<Field.Label for="subtitle">Subtitle</Field.Label>
<Input {...updateBookMetadata.fields.subtitle.as('text')} />
@@ -137,8 +141,14 @@
{/each}
</Field.Field>
<div class="grid grid-cols-[3fr_1fr] gap-2">
<!-- Series field -->
<Field.Field>
<Field.Label for="edition">Edition</Field.Label>
<Input {...updateBookMetadata.fields.edition.as('number')} />
{#each updateBookMetadata.fields.edition.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
<Field.Field>
<Field.Label for="series">Series</Field.Label>
<Input {...updateBookMetadata.fields.series.as('text')} />
@@ -147,18 +157,27 @@
{/each}
</Field.Field>
<!-- Series position field -->
<div class="grid grid-cols-2 gap-2">
<Field.Field>
<Field.Label for="series_position">Series position</Field.Label>
<Field.Label for="series_position">No.</Field.Label>
<Input {...updateBookMetadata.fields.series_position.as('text')} />
{#each updateBookMetadata.fields.series_position.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
<Field.Field>
<Field.Label for="language">Language</Field.Label>
<Input {...updateBookMetadata.fields.language.as('text')} />
{#each updateBookMetadata.fields.language.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
</div>
<!-- Authors field -->
<Field.Field>
{@render groupHeading('People and subjects')}
<Field.Field class="col-span-full">
<Field.Label for="authors">Authors</Field.Label>
<TagsInput
bind:value={authors}
@@ -174,8 +193,7 @@
{/each}
</Field.Field>
<!-- Tags field -->
<Field.Field>
<Field.Field class="col-span-full">
<Field.Label for="tags">Tags</Field.Label>
<TagsInput
bind:value={tags}
@@ -191,39 +209,8 @@
{/each}
</Field.Field>
<!-- Description field -->
<Field.Field>
<Field.Label for="description">Description</Field.Label>
<Textarea {...updateBookMetadata.fields.description.as('text')} />
{#each updateBookMetadata.fields.description.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
{@render groupHeading('Publication')}
<!-- Identifier fields -->
<Field.Field>
<Field.Label for="identifiers">Identifiers</Field.Label>
<div class="grid grid-cols-[5fr_10fr_0.5fr] gap-2">
{#each identifierKeys as _, idx}
<Input bind:value={identifierKeys[idx]} placeholder="Identifier..." />
<Input bind:value={identifierValues[idx]} placeholder="Value..." />
<Button variant="outline" size="icon" onclick={() => handleRemoveIdentifier(idx)}>
<Minus />
</Button>
{/each}
<Button variant="outline" onclick={() => handleAddIdentifier()}>
<Plus />
Add Identifier
</Button>
<input {...updateBookMetadata.fields.identifiers.as('text')} class="hidden" />
</div>
</Field.Field>
<!-- Publisher field -->
<div class="grid grid-cols-[2fr_1fr] gap-2">
<Field.Field>
<Field.Label for="publisher">Publisher</Field.Label>
<Input {...updateBookMetadata.fields.publisher.as('text')} />
@@ -232,18 +219,15 @@
{/each}
</Field.Field>
<!-- Published date field -->
<div class="grid grid-cols-2 gap-2">
<Field.Field>
<Field.Label for="published_date">Date published</Field.Label>
<Field.Label for="published_date">Published</Field.Label>
<Input {...updateBookMetadata.fields.published_date.as('date')} />
{#each updateBookMetadata.fields.published_date.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
</div>
<div class="grid grid-cols-[2fr_1fr_1fr] gap-2">
<!-- Pages field -->
<Field.Field>
<Field.Label for="pages">Pages</Field.Label>
<Input {...updateBookMetadata.fields.pages.as('number')} />
@@ -251,32 +235,44 @@
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
<!-- Language field -->
<Field.Field>
<Field.Label for="language">Language</Field.Label>
<Input {...updateBookMetadata.fields.language.as('text')} />
{#each updateBookMetadata.fields.language.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
<!-- Edition field -->
<Field.Field>
<Field.Label for="edition">Edition</Field.Label>
<Input {...updateBookMetadata.fields.edition.as('number')} />
{#each updateBookMetadata.fields.edition.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
</div>
</Field.Group>
</Field.Set>
</Card.Content>
<!-- Submit button -->
<Card.Footer class="flex-col gap-2 pt-6">
<Button type="submit" class="w-full">Save</Button>
</Card.Footer>
<Field.Field class="col-span-full">
<Field.Label for="identifiers">Identifiers</Field.Label>
<div class="flex flex-col gap-2">
{#each identifierKeys as _, idx}
<div class="grid grid-cols-[1fr_1.6fr_auto] gap-2">
<Input bind:value={identifierKeys[idx]} placeholder="ISBN, DOI…" />
<Input bind:value={identifierValues[idx]} placeholder="Value" />
<Button
type="button"
variant="outline"
size="icon"
title="Remove identifier"
onclick={() => handleRemoveIdentifier(idx)}
>
<Minus />
<span class="sr-only">Remove identifier</span>
</Button>
</div>
{/each}
<Button type="button" variant="outline" class="w-fit" onclick={handleAddIdentifier}>
<Plus />
Add identifier
</Button>
<input {...updateBookMetadata.fields.identifiers.as('text')} class="hidden" />
</div>
</Field.Field>
{@render groupHeading('Description')}
<Field.Field class="col-span-full">
<Field.Label for="description" class="sr-only">Description</Field.Label>
<Textarea {...updateBookMetadata.fields.description.as('text')} rows={6} />
{#each updateBookMetadata.fields.description.issues() ?? [] as issue}
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
</form>
</Card.Root>
@@ -5,13 +5,14 @@
import { Input } from '$lib/components/ui/input/index.js';
import { Textarea } from '$lib/components/ui/textarea/index';
import { Button } from '$lib/components/ui/button/index.js';
import { IconPicker } from '$lib/components/ui/icon-picker/index.js';
import { toast } from 'svelte-sonner';
import { goto, invalidate, invalidateAll } from '$app/navigation';
import { page } from '$app/state';
import { getLibraryState } from '$lib/state/library.svelte';
let { open = $bindable(false) } = $props();
let selectedIcon = $state('library');
const libraryState = getLibraryState();
</script>
@@ -35,6 +36,7 @@
const libraryName = createLibrary.fields.name.value();
form.reset();
selectedIcon = 'library';
open = false;
toast.success(`Library '${libraryName}' created.`);
@@ -48,7 +50,9 @@
>
<Field.Set>
<Field.Group>
<!-- Library name -->
<!-- Library name and Icon -->
<div class="flex items-end gap-3">
<div class="flex-1">
<Field.Field>
<Field.Label for="name">Library name</Field.Label>
<Input {...createLibrary.fields.name.as('text')} />
@@ -56,6 +60,14 @@
<Field.Error>{issue.message}</Field.Error>
{/each}
</Field.Field>
</div>
<div class="flex flex-col gap-2">
<Field.Label>Icon</Field.Label>
<IconPicker bind:value={selectedIcon} />
<input type="hidden" name="icon" value={selectedIcon} />
</div>
</div>
<!-- Description -->
<Field.Field>
@@ -0,0 +1,149 @@
import type { Book } from '$lib/schema';
/**
* What kind of thing a metadata field holds, which is the only thing that decides
* what you can do with it when two records disagree.
*/
export type FieldKind = 'text' | 'number' | 'date' | 'list' | 'keyed' | 'longtext';
/** An action offered on a field, beyond replacing it outright. */
export type FieldAction = 'replace' | 'merge' | 'append';
export interface FieldSpec {
/** The key sent in `BookMetadataUpdate`. */
key: string;
label: string;
kind: FieldKind;
group: string;
/** Extra actions past `replace`, which every field has. */
extra: FieldAction[];
}
/**
* The fields a merge can resolve, in the order and grouping the edit form uses.
*
* Deliberately a spec rather than markup: the merge workbench and, later, the
* provider review screen both render from this, so a field cannot exist in one and
* not the other. `cover` and `files` are absent because they are not choices
* files always come across and the cover has its own endpoint.
*/
export const MERGE_FIELDS: FieldSpec[] = [
{ key: 'title', label: 'Title', kind: 'text', group: 'Identity', extra: [] },
{ key: 'subtitle', label: 'Subtitle', kind: 'text', group: 'Identity', extra: [] },
{ key: 'edition', label: 'Edition', kind: 'number', group: 'Identity', extra: [] },
{ key: 'series', label: 'Series', kind: 'text', group: 'Identity', extra: [] },
{ key: 'series_position', label: 'No.', kind: 'text', group: 'Identity', extra: [] },
{ key: 'language', label: 'Language', kind: 'text', group: 'Identity', extra: [] },
// Order is meaningful for authors, so the second list is appended rather than
// interleaved; tags are a set, so they merge.
{
key: 'authors',
label: 'Authors',
kind: 'list',
group: 'People and subjects',
extra: ['append']
},
{ key: 'tags', label: 'Tags', kind: 'list', group: 'People and subjects', extra: ['merge'] },
{ key: 'publisher', label: 'Publisher', kind: 'text', group: 'Publication', extra: [] },
{ key: 'published_date', label: 'Published', kind: 'date', group: 'Publication', extra: [] },
{ key: 'pages', label: 'Pages', kind: 'number', group: 'Publication', extra: [] },
{
key: 'identifiers',
label: 'Identifiers',
kind: 'keyed',
group: 'Publication',
extra: ['merge']
},
{
key: 'description',
label: 'Description',
kind: 'longtext',
group: 'Description',
extra: ['append']
}
];
/** A field's value, in the shape `BookMetadataUpdate` expects to receive it. */
export type FieldValue = string | number | string[] | Record<string, string> | null;
/** Read one field off a book, flattening the relations the API returns as objects. */
export function readField(book: Book, key: string): FieldValue {
switch (key) {
case 'authors':
return book.authors.map((author) => author.name);
case 'tags':
return book.tags.map((tag) => tag.name);
case 'publisher':
return book.publisher?.name ?? null;
case 'series':
return book.series?.title ?? null;
case 'identifiers':
return book.identifiers ?? {};
default:
return (book as unknown as Record<string, FieldValue>)[key] ?? null;
}
}
/** Whether a field holds nothing, and so has no decision attached to it. */
export function isEmpty(value: FieldValue): boolean {
if (value === null || value === undefined || value === '') return true;
if (Array.isArray(value)) return value.length === 0;
if (typeof value === 'object') return Object.keys(value).length === 0;
return false;
}
/** Whether two field values say the same thing, order included for lists. */
export function isSame(left: FieldValue, right: FieldValue): boolean {
if (isEmpty(left) && isEmpty(right)) return true;
return JSON.stringify(left ?? null) === JSON.stringify(right ?? null);
}
/**
* Apply an action to a pair of values and return what the target becomes.
*
* `merge` on a keyed collection is per name and the target wins a clash, because
* `Identifier` is unique on `(name, book_id)` a book cannot hold both its print
* and its ebook ISBN, so the second one has nowhere to go.
*/
export function applyAction(
action: FieldAction,
kind: FieldKind,
target: FieldValue,
incoming: FieldValue
): FieldValue {
if (action === 'replace') return incoming;
if (kind === 'list') {
const current = Array.isArray(target) ? target : [];
const extra = Array.isArray(incoming) ? incoming : [];
// Order preserved, duplicates dropped — works for both append and merge.
return [...new Set([...current, ...extra])];
}
if (kind === 'keyed') {
return { ...(incoming as Record<string, string>), ...(target as Record<string, string>) };
}
if (kind === 'longtext') {
const current = typeof target === 'string' ? target.trim() : '';
const extra = typeof incoming === 'string' ? incoming.trim() : '';
return [current, extra].filter(Boolean).join('\n\n');
}
return incoming;
}
/** How a value reads in the inert reference column. */
export function displayValue(value: FieldValue): string {
if (isEmpty(value)) return '';
if (Array.isArray(value)) return value.join(', ');
if (typeof value === 'object') {
return Object.entries(value as Record<string, string>)
.map(([name, id]) => `${name}: ${id}`)
.join(' · ');
}
return String(value);
}
@@ -0,0 +1,394 @@
<script lang="ts">
import { untrack } from 'svelte';
import { ArrowRight, GitMerge, Plus, Undo2 } from '@lucide/svelte';
import * as Dialog from '$lib/components/ui/dialog/index.js';
import { Button } from '$lib/components/ui/button/index.js';
import { Input } from '$lib/components/ui/input/index.js';
import { Textarea } from '$lib/components/ui/textarea/index.js';
import { Badge } from '$lib/components/ui/badge/index.js';
import BookImage from '$lib/components/view/book-image.svelte';
import GeneratedCover from '$lib/components/view/generated-cover.svelte';
import type { Book } from '$lib/schema';
import { mergeInto } from './merge';
import {
MERGE_FIELDS,
applyAction,
displayValue,
isEmpty,
isSame,
readField,
type FieldAction,
type FieldSpec,
type FieldValue
} from './field-spec';
let {
books,
libraryId,
open = $bindable(),
onmerged
}: {
books: Book[];
libraryId: number | string;
open: boolean;
/** Given the record that survived and the ones folded into it and deleted. */
onmerged?: (survivor: Book, folded: Book[]) => void;
} = $props();
// Seeded once per mount. The dialog is keyed on the group upstream, so a
// different group gets a fresh workbench rather than the previous one's draft.
let survivorId = $state(untrack(() => books[0]?.id));
let candidateId = $state(untrack(() => books[1]?.id));
let draft = $state<Record<string, FieldValue>>({});
let busy = $state(false);
const survivor = $derived(books.find((book) => book.id === survivorId) ?? books[0]);
const candidates = $derived(books.filter((book) => book.id !== survivorId));
/** The records that will be deleted — the same set, named for what happens to them. */
const folded = $derived(candidates);
const candidate = $derived(candidates.find((book) => book.id === candidateId) ?? candidates[0]);
/** The survivor's stored value for a field, or the draft if it has been touched. */
function current(field: FieldSpec): FieldValue {
return field.key in draft ? draft[field.key] : readField(survivor, field.key);
}
function take(field: FieldSpec, action: FieldAction) {
draft[field.key] = applyAction(
action,
field.kind,
current(field),
readField(candidate, field.key)
);
}
function undo(field: FieldSpec) {
delete draft[field.key];
draft = { ...draft };
}
/**
* Fill only the fields the survivor has nothing in.
*
* The safe bulk action, and the one worth reaching for: it cannot overwrite a
* value, so it needs no per-field protection to be pressed without reading.
*/
function fillEmpty() {
for (const field of MERGE_FIELDS) {
const incoming = readField(candidate, field.key);
if (isEmpty(current(field)) && !isEmpty(incoming)) draft[field.key] = incoming;
}
}
const changed = $derived(Object.keys(draft));
const differing = $derived(
candidate
? MERGE_FIELDS.filter((field) => !isSame(current(field), readField(candidate, field.key)))
: []
);
/** Fields shown as a row: anything the two disagree on, plus anything edited. */
const shown = $derived(
MERGE_FIELDS.filter((field) => differing.includes(field) || field.key in draft)
);
const agreed = $derived(MERGE_FIELDS.filter((field) => !shown.includes(field)));
const groups = $derived([...new Set(shown.map((field) => field.group))]);
async function submit() {
if (busy) return;
busy = true;
const merged = await mergeInto(libraryId, survivor, folded, changed.length ? draft : undefined);
busy = false;
if (merged) {
open = false;
onmerged?.(survivor, folded);
}
}
</script>
{#snippet cover(book: Book, size: string)}
<span class="{size} shrink-0 overflow-hidden rounded-sm border bg-muted">
{#if book.cover_image}
<BookImage src="/api/{book.cover_image}" class="h-full w-full object-cover" />
{:else}
<!-- Drawn rather than left blank, so the rail tells two coverless books
apart the same way the shelves do. -->
<GeneratedCover {book} />
{/if}
</span>
{/snippet}
<!-- The inert reference column: the same shape as the control opposite it, with
nothing that invites a click. -->
{#snippet reference(field: FieldSpec, book: Book)}
{@const value = readField(book, field.key)}
<div class="flex min-w-0 flex-col gap-1">
<span class="text-[11px] text-muted-foreground">{field.label}</span>
{#if isEmpty(value)}
<span class="min-h-8 py-1 text-sm text-muted-foreground italic">empty</span>
{:else if field.kind === 'list'}
<span class="flex min-h-8 flex-wrap items-center gap-1 py-0.5">
{#each value as string[] as item (item)}
<Badge variant="secondary" class="font-normal">{item}</Badge>
{/each}
</span>
{:else}
<span class="min-h-8 py-1 text-sm break-words">{displayValue(value)}</span>
{/if}
</div>
{/snippet}
<Dialog.Root bind:open>
<Dialog.Content
class="flex h-[85vh] max-w-[min(72rem,95vw)] flex-col gap-0 overflow-hidden p-0 sm:max-w-[min(72rem,95vw)]"
>
<Dialog.Header class="shrink-0 space-y-0 border-b px-5 py-3 text-left">
<Dialog.Title class="font-serif text-base font-normal">
Merge {books.length} books
</Dialog.Title>
<Dialog.Description class="text-xs">
{candidates.length}
{candidates.length === 1 ? 'record is' : 'records are'} deleted. Their files move onto the book
you keep — nothing is removed from disk.
</Dialog.Description>
</Dialog.Header>
<div class="grid min-h-0 flex-1 grid-cols-1 md:grid-cols-[15rem_1fr]">
<!-- Rail: the books being folded in, one open at a time -->
<aside
class="flex min-w-0 flex-col overflow-y-auto border-b bg-sidebar md:border-r md:border-b-0"
>
<p
class="px-3 pt-3 pb-1 font-mono text-[10px] tracking-widest text-muted-foreground uppercase"
>
Taking from · {candidates.length}
</p>
{#each candidates as book (book.id)}
{@const count = MERGE_FIELDS.filter(
(field) => !isSame(readField(survivor, field.key), readField(book, field.key))
).length}
<button
type="button"
onclick={() => (candidateId = book.id)}
class="flex items-start gap-2 border-l-2 px-3 py-2 text-left transition-colors hover:bg-muted/50 {book.id ===
candidate?.id
? 'border-l-primary bg-background'
: 'border-l-transparent'}"
>
{@render cover(book, 'h-10 w-7')}
<span class="min-w-0 flex-1">
<span class="line-clamp-2 font-serif text-[13px]">{book.title}</span>
<span class="block text-[10px] text-muted-foreground">
#{book.id} · {book.files.length}
{book.files.length === 1 ? 'file' : 'files'}
</span>
{#if count === 0}
<Badge variant="secondary" class="mt-1 text-[10px] font-normal">
nothing to take
</Badge>
{/if}
</span>
</button>
{/each}
</aside>
<div class="flex min-h-0 flex-col">
<!-- Which record survives -->
<div class="flex shrink-0 flex-wrap items-center gap-2 border-b px-5 py-2.5">
<span class="text-[11px] text-muted-foreground">Keeping</span>
{#each books as book (book.id)}
<button
type="button"
onclick={() => (survivorId = book.id)}
class="flex items-center gap-1.5 rounded-md border px-2 py-1 text-xs transition-colors {book.id ===
survivorId
? 'border-primary bg-accent text-accent-foreground'
: 'text-muted-foreground hover:bg-muted'}"
>
{@render cover(book, 'h-6 w-4')}
<span class="max-w-32 truncate">#{book.id}</span>
</button>
{/each}
<div class="ml-auto flex gap-2">
<Button variant="outline" size="sm" onclick={fillEmpty} disabled={!candidate}>
Fill empty fields
</Button>
</div>
</div>
<div class="min-h-0 flex-1 overflow-y-auto px-5 py-4">
{#if !candidate}
<p class="text-sm text-muted-foreground">Nothing left to fold in.</p>
{:else if shown.length === 0}
<p class="text-sm text-muted-foreground">
These records agree on every field. Merging keeps
<span class="font-medium text-foreground">#{survivor.id}</span> and moves the others' files
onto it.
</p>
{:else}
{#each groups as group (group)}
<h3
class="mt-5 border-b pb-1 font-mono text-[10px] tracking-widest text-primary uppercase first:mt-0"
>
{group}
</h3>
{#each shown.filter((field) => field.group === group) as field (field.key)}
{@const value = current(field)}
{@const incoming = readField(candidate, field.key)}
{@const edited = field.key in draft}
<div class="grid grid-cols-[1fr_auto_1fr] items-center gap-3 border-b py-2">
{@render reference(field, candidate)}
<!-- Actions, on the row rather than stacked beside it -->
<div class="flex items-center gap-1">
{#if edited}
<Button
variant="ghost"
size="icon"
class="size-7"
title="Undo"
onclick={() => undo(field)}
>
<Undo2 class="size-3.5" />
<span class="sr-only">Undo {field.label}</span>
</Button>
{/if}
{#if !isSame(value, incoming)}
{#if isEmpty(value)}
<!-- Nothing to weigh, so it is an offer rather than a choice -->
<Button
size="sm"
class="h-7 gap-1 px-2 text-[11px]"
onclick={() => take(field, 'replace')}
>
<ArrowRight class="size-3.5" />
Take
</Button>
{:else}
<Button
variant="outline"
size="icon"
class="size-7"
title="Replace"
onclick={() => take(field, 'replace')}
>
<ArrowRight class="size-3.5" />
<span class="sr-only">Replace {field.label}</span>
</Button>
{#each field.extra as action (action)}
<Button
variant="outline"
size="icon"
class="size-7"
title={action === 'merge' ? 'Merge' : 'Append'}
onclick={() => take(field, action)}
>
{#if action === 'merge'}
<GitMerge class="size-3.5" />
{:else}
<Plus class="size-3.5" />
{/if}
<span class="sr-only">{action} {field.label}</span>
</Button>
{/each}
{/if}
{/if}
</div>
<!-- The survivor's side: the edit form, live -->
<div class="flex min-w-0 flex-col gap-1">
<span class="flex items-center gap-1.5 text-[11px] text-muted-foreground">
{field.label}
{#if edited}
<Badge class="h-4 px-1.5 text-[9px] font-semibold">taken</Badge>
{/if}
</span>
{#if field.kind === 'longtext'}
<Textarea
rows={4}
class="text-sm"
value={(value as string) ?? ''}
oninput={(event) => (draft[field.key] = event.currentTarget.value)}
/>
{:else if field.kind === 'list' || field.kind === 'keyed'}
<!-- Edited through the take actions; typing here would need the
tags and key/value editors, which belong to the edit form. -->
<div
class="flex min-h-8 flex-wrap items-center gap-1 rounded-md border bg-background px-2 py-1"
>
{#if isEmpty(value)}
<span class="text-sm text-muted-foreground"></span>
{:else if field.kind === 'list'}
{#each value as string[] as item (item)}
<Badge variant="secondary" class="font-normal">{item}</Badge>
{/each}
{:else}
{#each Object.entries(value as Record<string, string>) as [name, id] (name)}
<Badge variant="secondary" class="font-mono text-[10px] font-normal">
{name}: {id}
</Badge>
{/each}
{/if}
</div>
{:else}
<Input
type={field.kind === 'number'
? 'number'
: field.kind === 'date'
? 'date'
: 'text'}
class="h-8 text-sm"
value={(value as string | number) ?? ''}
oninput={(event) =>
(draft[field.key] =
field.kind === 'number'
? Number(event.currentTarget.value) || null
: event.currentTarget.value)}
/>
{/if}
</div>
</div>
{/each}
{/each}
{#if agreed.length > 0}
<p class="pt-3 text-center text-[11px] text-muted-foreground italic">
{agreed.map((field) => field.label).join(', ')} — identical in both
</p>
{/if}
{/if}
</div>
</div>
</div>
<Dialog.Footer class="shrink-0 items-center gap-2 border-t px-5 py-3 sm:justify-between">
<p class="text-xs text-muted-foreground">
{#if changed.length}
{changed.length}
{changed.length === 1 ? 'change' : 'changes'} pending · this cannot be undone
{:else}
Metadata is left as #{survivor.id} has it · this cannot be undone
{/if}
</p>
<div class="flex gap-2">
<Button variant="outline" onclick={() => (open = false)}>Cancel</Button>
<Button onclick={submit} disabled={busy || candidates.length === 0}>
Merge into “{survivor.title}
</Button>
</div>
</Dialog.Footer>
</Dialog.Content>
</Dialog.Root>
@@ -0,0 +1,44 @@
import { invalidate } from '$app/navigation';
import { toast } from 'svelte-sonner';
import { mergeBooks } from '$lib/api/book.remote';
import type { Book } from '$lib/schema';
import type { FieldValue } from './field-spec';
/**
* Fold `folded` into `survivor`, and tell the reader how it went.
*
* Shared because a merge is reachable two ways the workbench, where the reader
* resolved the metadata field by field, and the quick action beside it, which
* takes the survivor's metadata as it stands. Only the `metadata` argument
* differs, and the invalidation and the wording should not.
*
* @returns whether the merge went through; the caller decides what to close or
* clear, and has no toast of its own to write either way.
*/
export async function mergeInto(
libraryId: number | string,
survivor: Book,
folded: Book[],
metadata?: Record<string, FieldValue>
): Promise<boolean> {
try {
await mergeBooks({
library_id: libraryId,
survivor_id: survivor.id,
merged_ids: folded.map((book) => book.id),
metadata
});
// Both, because a merge is started from the duplicates review and from the
// library's selection toolbar, and each page depends on a different one.
await Promise.all([invalidate('app:books'), invalidate('app:duplicate-books')]);
toast.success(`Merged into “${survivor.title}`);
return true;
} catch (error) {
console.error('Failed to merge books', error);
toast.error(error instanceof Error ? error.message : 'Could not merge these books');
return false;
}
}

Some files were not shown because too many files have changed in this diff Show More