The hash lives in the database and the file does not, so a file deleted behind the app's back went on refusing its own replacement. Matches are now checked against disk, and re-adding a book's own missing file writes it back into the row that already describes it.
204 lines
12 KiB
Markdown
204 lines
12 KiB
Markdown
# Chitai backend
|
|
|
|
Litestar REST API for the eBook library. See the repo-root `AGENTS.md` for the overall picture and
|
|
dev-environment setup.
|
|
|
|
**Stack:** Python 3.13 · Litestar 2 · advanced-alchemy over async SQLAlchemy 2 · asyncpg ·
|
|
PostgreSQL 17 · Alembic (through advanced-alchemy's `alchemy` CLI) · pydantic-settings · uv.
|
|
|
|
## Layering
|
|
|
|
```
|
|
controllers/ HTTP surface only — parse, delegate, serialise. Keep thin.
|
|
services/ Business logic. One SQLAlchemyAsyncRepositoryService subclass per aggregate.
|
|
database/models/ SQLAlchemy models.
|
|
schemas/ Pydantic DTOs for request bodies and responses.
|
|
```
|
|
|
|
Controllers call **service** methods, not the repository. `app.py` assembles the app: route
|
|
handlers, JWT auth, exception handlers, the SQLAlchemy plugin, and two lifespan context managers.
|
|
|
|
## advanced-alchemy idioms
|
|
|
|
These are the conventions that are easy to get wrong if you write plain SQLAlchemy here:
|
|
|
|
- Models extend `BigIntAuditBase` (adds id/created_at/updated_at); pure link and child rows extend
|
|
`BigIntBase` (e.g. `Identifier`, `FileMetadata`, `BookAuthorLink`).
|
|
- A service declares an inner repository and points at it:
|
|
|
|
```python
|
|
class BookService(SQLAlchemyAsyncRepositoryService[Book]):
|
|
class Repo(SQLAlchemyAsyncRepository[Book]):
|
|
model_type = Book
|
|
repository_type = Repo
|
|
```
|
|
|
|
- Transform incoming data with the **`to_model_on_create` / `to_model_on_update` hooks**, not by
|
|
overriding `create`/`update` wholesale — see `services/book.py:407` onward.
|
|
- Serialise responses through `service.to_schema(obj, schema_type=s.SomeRead)`; for lists,
|
|
`to_schema(items, total, filters, schema_type=…)` produces the `OffsetPagination` envelope.
|
|
- `Author`, `Tag`, `Publisher` and `BookSeries` are deduplicated with **`as_unique_async`**. Never
|
|
construct them directly when attaching to a book — use
|
|
`await Author.as_unique_async(session, name=name)` as
|
|
`BookService._populate_with_unique_relationships` does, or you will create duplicate rows.
|
|
- Domain methods on services are named for the domain (`create_book`, `update_book`, `add_files`),
|
|
deliberately distinct from the inherited CRUD names.
|
|
|
|
## Dependency injection
|
|
|
|
`services/dependencies.py` is the hub; controllers wire providers in their `dependencies` dict.
|
|
|
|
- Most providers come from `create_service_provider(SomeService, …)` — one line each.
|
|
- `provide_book_service` is hand-written because it must inject eager loads *and* scope
|
|
user-specific rows: `selectinload` for authors/tags/files/etc., plus `with_loader_criteria` so
|
|
`BookProgress` and `BookListLink` only load rows belonging to `current_user`. If you add a
|
|
relationship that the API returns, add it to that `load` list.
|
|
- `create_book_filter_dependencies` intentionally **overrides** advanced-alchemy's stock providers:
|
|
the search filter becomes a trigram search, and order-by gains a `random` sort order. Do not
|
|
replace it with the stock `create_filter_dependencies`.
|
|
- `get_library_by_id` resolves the target library from either a `library_id` query param or the
|
|
book's own `library_id`, and raises 404 for either miss.
|
|
|
|
## Filters
|
|
|
|
`services/filters/` holds `StatementFilter` dataclasses that compose into any `list` /
|
|
`list_and_count` call: `TagFilter`, `AuthorFilter`, `BookshelfFilter`, `ProgressFilter`,
|
|
`TrigramSearchFilter`, `CustomOrderBy`, `FileHashFilter`, plus the `*LibraryFilter` variants used by
|
|
OPDS.
|
|
|
|
To add list-filtering behaviour: write the dataclass here, add a `provide_*_filter` function in
|
|
`dependencies.py`, register it in the controller's `dependencies`, and fold it into
|
|
`provide_book_filters` — the controller handler itself does not change.
|
|
|
|
## Authentication — three schemes
|
|
|
|
| Scheme | Where | Used by |
|
|
| --- | --- | --- |
|
|
| JWT bearer (`OAuth2PasswordBearerAuth`) | `app.py` | The web frontend and the main API. Public paths are listed in its `exclude=[…]`. |
|
|
| HTTP Basic (`middleware/basic_auth.py`) | `OpdsController` | E-reader / OPDS clients, which only speak Basic. |
|
|
| `X-AUTH-USER` API key (`middleware/kosync_auth.py`) | `KosyncController` | KOReader devices; the key maps to a `KosyncDevice` row, which maps to a user. |
|
|
|
|
Each middleware resolves a `User` onto the connection; the matching
|
|
`provide_user_via_basic_auth` / `provide_user_via_kosync_auth` dependencies expose it to handlers.
|
|
|
|
## Filesystem behaviour
|
|
|
|
The backend owns files on disk, not just rows:
|
|
|
|
- **Layout** — `services/filesystem_library.py` (`BookPathGenerator`) renders a Jinja2 template
|
|
against book metadata to decide where a book lives under the library's `root_path`
|
|
(default: `author/series/position - title/`).
|
|
- **One directory per book, never shared.** The generated path is a pure function of the metadata,
|
|
so two books with the same author and title produce the same one — two editions, or an
|
|
`allow_duplicates` copy. `BookService._reserve_book_path` moves the later one to `title (2)`
|
|
before anything is written, and `update_book` reserves the same way so a rename cannot move a
|
|
book in on top of another. This matters because `book.path` is what deletes, moves and file
|
|
lookups act on: books sharing a directory means one overwrites the other's files, and deleting
|
|
either takes both. `_unused_path` does the same job for filenames within a directory.
|
|
A book that already has a `path` keeps it — `add_files` must follow the book, not the template.
|
|
- **Metadata extraction** — `services/metadata_extractor.py` reads EPUB (ebooklib) and PDF
|
|
(pypdfium2) files; extracted values fill only *empty* fields on the incoming payload.
|
|
- **Covers** — converted to WebP with a UUID filename under `settings.book_cover_path`, served by a
|
|
static-files router mounted at `/covers`.
|
|
- **Consume directory** — `services/consume.py` (`ConsumeDirectoryWatcher`) watches
|
|
`settings.consume_path` with `watchfiles`, creates one subdirectory per library slug, batches
|
|
additions (3 s debounce) and imports them via `BookService.create_many_from_existing_files`.
|
|
Started as an asyncio task from the `setup_directory_watcher` lifespan hook.
|
|
- **Updates move files.** `BookService.update_book` regenerates the path from the new metadata and,
|
|
if it differs, moves the directory contents and prunes empty parents. Keep that in mind before
|
|
changing metadata handling.
|
|
|
|
## Duplicate detection
|
|
|
|
Every ingest path screens incoming files against what is already stored, keyed on
|
|
**`(hash, size)`** — never the hash alone, because it samples 12 KiB (see below) and
|
|
EPUBs from one toolchain often share their first window. `FileMetadata.hash` carries a
|
|
plain, deliberately **non-unique** index: a collision must not be able to fail an import,
|
|
and older databases may already hold duplicates.
|
|
|
|
Scope comes from `CHITAI_DUPLICATE_SCOPE` (`library`, the default | `global` | `off`).
|
|
|
|
The policy differs by how deliberate the import is:
|
|
|
|
| Path | Behaviour |
|
|
| --- | --- |
|
|
| `create_many_from_files` (browser bulk) | Skip per file, skip a whole group whose files are all known, report everything skipped in `ImportResult.duplicates`. Re-dropping a folder to pick up what is new is the case this serves. |
|
|
| `create_book` (single, with metadata) | All-or-nothing: raises `DuplicateFilesError`, which `controllers/book.py` renders as a **409** carrying the refused files in `extra`. |
|
|
| `add_files` | A file the book already carries is a no-op; one stored under another book raises `DuplicateFilesError`. |
|
|
| `create_many_from_existing_files` (consume watcher) | Skips, and **moves the file to `CHITAI_DUPLICATE_PATH/<library slug>/`** — nothing is deleted, and it cannot stay put because `watchfiles` only reports additions. That path must stay outside `consume_path` or the watcher re-imports it and tries to read the directory name as a library slug. |
|
|
|
|
`allow_duplicates=true` overrides all of it, on every endpoint. Keep that working — the
|
|
hash is not proof of identity, so a wrong verdict has to be recoverable, and a scripted
|
|
import needs a way through. The **web UI deliberately does not offer it**: storing the
|
|
same bytes twice splits reading progress and shelf membership across two records that
|
|
can never converge, which is nothing anyone wants on purpose.
|
|
|
|
Three things to preserve when touching this code:
|
|
|
|
- **A match only counts while the file is on disk.** `find_duplicate_files` stats each
|
|
candidate, and `add_files` writes a missing file back into the row that already
|
|
describes it (`_restore_file`) instead of adding a second row beside it. The hash
|
|
lives in the database and the file does not, so without this a file deleted behind
|
|
the app's back would go on refusing its own replacement.
|
|
|
|
- **Screening runs before anything is written.** `fingerprint_upload` reads the spooled
|
|
upload and rewinds it; the resulting fingerprints are handed to `_save_book_files`,
|
|
which skips its own `StreamingHasher` when it already has the answer. Passing them
|
|
through is what keeps the file from being read twice.
|
|
- **`_screen_for_duplicates` extends the `known` dict as it goes**, so the same bytes
|
|
submitted twice in one request are caught. Those duplicates report `book_id: None` —
|
|
there is no row to point at yet.
|
|
|
|
`POST /books/duplicates` answers the same question from fingerprints alone, for clients
|
|
that want to ask before uploading anything.
|
|
|
|
## KOReader hashing
|
|
|
|
`services/utils.py` reimplements KOReader's partial-MD5 document identifier: 1 KiB samples at
|
|
offsets produced by LuaJIT's 32-bit `bit.lshift`, including its shift-masking wrap-around
|
|
(`shift & 0x1F`, so `i=-1` yields offset 0). `_lshift32` looks wrong and is not — the overflow is
|
|
what makes hashes match real devices. Don't "simplify" it; `tests/integration/test_file_hash.py`
|
|
guards the behaviour.
|
|
|
|
## Database
|
|
|
|
- **`pg_trgm` is required.** `Book.__table_args__` declares a GIN trigram index on `title`, which
|
|
`TrigramSearchFilter` uses for fuzzy title search. The extension is enabled by a migration and, in
|
|
tests, by `conftest.py`.
|
|
- Migrations live in `migrations/versions/`. Generate with
|
|
`alchemy --config chitai.database.config.config make-migrations` (add `--no-autogenerate` for a
|
|
blank revision), apply with `… upgrade`. `database/config.py` sets `create_all=False`, so nothing
|
|
is auto-created at runtime; production applies migrations from `entrypoint.sh`.
|
|
- Sessions use `expire_on_commit=False` and Litestar's `before_send_handler="autocommit"`, so a
|
|
handler that returns 2xx commits automatically.
|
|
|
|
## Testing
|
|
|
|
`pytest tests/` — `asyncio_mode = "auto"`, so async tests need no marker.
|
|
|
|
- `tests/unit/` exercises services directly against a real session; `tests/integration/` drives the
|
|
whole app through `AsyncTestClient(app=create_app())`.
|
|
- The database is a throwaway container from `pytest-databases`, so **Docker must be running**.
|
|
- `tests/conftest.py` provides the shared fixtures: `client`, `authenticated_client`,
|
|
`other_authenticated_client` (a second user, for access-control tests), one fixture per service,
|
|
`test_user` / `test_library`, and an autouse fixture that redirects cover storage into `tmp_path`.
|
|
- `tests/integration/conftest.py` monkeypatches the module-level alchemy `config` onto the test
|
|
engine/sessionmaker and drops+recreates+reseeds the schema for every test.
|
|
- Real EPUB and PDF fixtures live in `tests/data_files/`.
|
|
|
|
## Known rough edges
|
|
|
|
Observed in the current tree — don't mistake these for intentional patterns to copy:
|
|
|
|
- `controllers/book.py` — file-level TODO: `book_id` is a path parameter on some endpoints and a
|
|
query parameter on others. `set_book_progress_batch` does a documented N+1 (one select + one
|
|
upsert per book).
|
|
- `services/filesystem_library.py` — TODO to replace Jinja2 templating with simple placeholders;
|
|
`generate_filename` accepts a `filename_template` but currently ignores it and returns the
|
|
original filename.
|
|
- `app.py` — `watcher_task` is declared as a module-level global but assigned locally inside
|
|
`setup_directory_watcher`, so the global is never populated (cancellation still works via the
|
|
closure).
|
|
- `BookService.get_files` (the multi-book ZIP download) opens `Path(file.path)`, but `file.path` is
|
|
stored relative to `book.path` — worth verifying before relying on that endpoint.
|