# TODO Known issues and deferred work. Agent-facing notes belong in the `AGENTS.md` files; this is for things that are broken or missing and not yet scheduled. ## Backend ### Identifier extraction drops most ISBNs `backend/src/chitai/services/metadata_extractor.py` — `EpubExtractor._extract_identifiers` The EPUB path validates `DC:identifier` values verbatim: ```python for id in epub.get_metadata("DC", "identifier"): if is_valid_isbn(id[0]): ... ``` `is_valid_isbn` branches on `len(isbn)` being exactly 10 or 13, so anything carrying formatting fails. Two consequences: - **Hyphenated ISBNs are silently dropped.** `978-0-486-28211-4` is 17 characters, so it never reaches the checksum. Most EPUBs write ISBNs hyphenated, so the majority are lost. The PDF path already does this correctly — `_extract_isbns` calls `match.replace("-", "")` before validating. - **`urn:isbn:` prefixes are dropped** for the same reason. This is a common EPUB form. Fix: normalise before validating — strip a leading `urn:isbn:`, then remove everything that isn't `0-9` or `X`. Reuse the PDF path's approach rather than duplicating it. Note this only runs at upload, so fixing it changes nothing for books already imported. A backfill would need to re-read the files on disk. ### Non-ISBN identifiers are discarded Same function. Anything that isn't a valid ISBN is thrown away, including values EPUBs routinely carry: `urn:uuid:…`, `calibre:…`, Google Books volume IDs and ASINs. The `Identifier` model is already generic (`name` + `value`, unique per book), so storing them needs no schema change — only the extractor decides what survives. The frontend already renders ISBN, ASIN and DOI as links and shows unknown types as plain values, so anything stored will display sensibly. Worth adding at the same time: - A DOI regex (`10.\d{4,9}/\S+`) alongside the ISBN scan in `PdfExtractor._extract_isbns` — academic PDFs carry one and it is the most useful identifier they have. - Deriving ISBN-10 from ISBN-13 when only the latter is present. It is a pure checksum conversion and doubles the chance of an external lookup matching. ## Frontend ### No CSP, so scripted EPUBs run against the app origin EPUB files may contain JavaScript. foliate-js renders each section in an iframe from a **same-origin** `blob:` URL and cannot sandbox it — `allow-scripts` is required, and blob URLs inherit the embedder's origin — so script inside a book can reach `/api/*` with the session cookie attached. foliate's own README says not to use it without a Content Security Policy blocking scripts. The obvious policy is `kit.csp` in `svelte.config.js` with `script-src: ['self']`, and deliberately no `default-src` (it would also cover `style-src`/`img-src`/`font-src` and kill both the book's own blob: assets and the inline `