From d8f8a0ab9577ab7907acca5891c42e9c53008431 Mon Sep 17 00:00:00 2001 From: patrick Date: Tue, 11 Aug 2026 20:08:16 -0400 Subject: [PATCH] docs: add TODO.md --- TODO.md | 84 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 84 insertions(+) create mode 100644 TODO.md diff --git a/TODO.md b/TODO.md new file mode 100644 index 0000000..7989dea --- /dev/null +++ b/TODO.md @@ -0,0 +1,84 @@ +# TODO + +Known issues and deferred work. Agent-facing notes belong in the `AGENTS.md` files; +this is for things that are broken or missing and not yet scheduled. + +## Backend + +### Identifier extraction drops most ISBNs + +`backend/src/chitai/services/metadata_extractor.py` — `EpubExtractor._extract_identifiers` + +The EPUB path validates `DC:identifier` values verbatim: + +```python +for id in epub.get_metadata("DC", "identifier"): + if is_valid_isbn(id[0]): + ... +``` + +`is_valid_isbn` branches on `len(isbn)` being exactly 10 or 13, so anything carrying +formatting fails. Two consequences: + +- **Hyphenated ISBNs are silently dropped.** `978-0-486-28211-4` is 17 characters, so it + never reaches the checksum. Most EPUBs write ISBNs hyphenated, so the majority are lost. + The PDF path already does this correctly — `_extract_isbns` calls + `match.replace("-", "")` before validating. +- **`urn:isbn:` prefixes are dropped** for the same reason. This is a common EPUB form. + +Fix: normalise before validating — strip a leading `urn:isbn:`, then remove everything +that isn't `0-9` or `X`. Reuse the PDF path's approach rather than duplicating it. + +Note this only runs at upload, so fixing it changes nothing for books already imported. +A backfill would need to re-read the files on disk. + +### Non-ISBN identifiers are discarded + +Same function. Anything that isn't a valid ISBN is thrown away, including values EPUBs +routinely carry: `urn:uuid:…`, `calibre:…`, Google Books volume IDs and ASINs. + +The `Identifier` model is already generic (`name` + `value`, unique per book), so storing +them needs no schema change — only the extractor decides what survives. The frontend +already renders ISBN, ASIN and DOI as links and shows unknown types as plain values, so +anything stored will display sensibly. + +Worth adding at the same time: + +- A DOI regex (`10.\d{4,9}/\S+`) alongside the ISBN scan in `PdfExtractor._extract_isbns` + — academic PDFs carry one and it is the most useful identifier they have. +- Deriving ISBN-10 from ISBN-13 when only the latter is present. It is a pure checksum + conversion and doubles the chance of an external lookup matching. + +## Frontend + +### Cover dimensions are unknown until load + +`book-cover.svelte` renders covers at a fixed height with natural width so nothing is +cropped or distorted. Because the intrinsic size is not known ahead of time, the text +beside the cover settles once when the image loads. `min-width` bounds the movement but +does not remove it. + +Removing it properly means storing cover dimensions at ingest and emitting them as +`width`/`height` attributes so the browser reserves exact space. That is a model change +plus a migration. + +### No series navigation + +`Book` has `series` and `series_position`, and the detail page shows both, but there is no +way to reach the other volumes. The books list endpoint filters by author, publisher, tag, +shelf and progress — a `series` filter would need adding in +`backend/src/chitai/services/filters/book.py` and wiring through +`services/dependencies.py`, following the existing `AuthorFilter` pattern. + +### Book pages are sparse for most books + +EPUB metadata is thin: most imports arrive with a title, an author and nothing else. Two +independent directions, neither started: + +- **Use what exists.** "More by this author" (the list endpoint already accepts + `authors=`), and exposing `Book.created_at` and `BookProgress.updated_at` — both are in + the database, neither is on `BookRead` / `BookProgressRead`. +- **Fetch from outside.** Open Library or Google Books lookup by ISBN for descriptions and + covers. This is what actually fixes the sparseness, but it needs outbound requests, rate + limiting, a manual-vs-automatic decision, and a rule for not clobbering hand-edited + metadata.