Skip to content

fix(ci): mirror to Codeberg only from canonical repo, not forks - #124

Closed
alphatownsman wants to merge 2423 commits into
doubaniux:masterfrom
neodb-social:fix/mirror-canonical-repo
Closed

alphatownsman wants to merge 2423 commits into
doubaniux:masterfrom
neodb-social:fix/mirror-canonical-repo

Conversation

@alphatownsman

Copy link
Copy Markdown
Contributor

Summary

  • Tighten the Codeberg mirror workflow guard from github.repository_owner == 'neodb-social' to github.repository == 'neodb-social/neodb', so the mirror job runs only on the canonical repo and never on forks or other repos under the org.

Test plan

  • pre-commit YAML checks pass
  • Verify the Mirror to Codeberg job is skipped on forks after merge

Your Name added 30 commits April 19, 2026 02:07
Address PR #1467 review feedback:
  - Remove sync_credits_from_metadata() call on edit-page GET so the
    view stays side-effect free and does not bypass the is_editable
    permission check (which is only enforced for POST).
  - Move display canonicalization into EditForm: when an ItemCredit
    already links a name to a People, swap the form's initial value
    for /person|organization/<uuid>. Pure read, no DB writes, same
    UX (indicator + name chip) as before.
  - In sync_credits_from_metadata, fetch the item's managed-role
    credits in a single query and resolve every /people|person
    |organization/<uuid> reference with one bulk People query, so
    items with many roles or many linked credits no longer issue an
    O(roles*people) query pattern.
Mirror the paste-URL behavior of the general search view: if the query
on the people/org search page contains "://", try to resolve it as a
local site redirect, a known external site (fetch + redirect to item),
a safe external URL when r=1, or fall back to the internet-wide fetch.
Addresses review comment on #1473: the URL-paste branch was duplicated
between search and people_search. Extract resolve_url_query() and have
both call sites delegate to it so future changes to URL resolution stay
in one place. Restores the "skip detecting redirection to avoid timeout"
comment that was dropped in the duplicated block.
On `/person/<id>` and `/organization/<id>` the works list showed both
the parent and its children (Work + Edition, TVShow + TVSeason, etc.)
whenever the person was credited on both, which cluttered the page.

Hide a descendant when its parent (or TVShow grandparent for a
TVEpisode) is also credited for the same person+role. The filter is a
handful of indexed id-only lookups keyed on the already-materialized
`item_ids`, so no additional item rows are fetched; pagination and
total counts both reflect the filtered list.
Review feedback (#1476): the descendant-hiding pass ran against the raw
`ItemPeopleRelation` id list, so:

- a soft-deleted or merged parent still masked its active child, making
  both disappear from the page;
- the shelf status filter could drop the parent while the pre-filter
  exclusion had already hidden the child, leaving an empty page.

Compute `hidden_ids` from `works_qs.values_list("pk")` after the active-
item and status filters have been applied, so only ancestors that will
actually render hide their descendants. Adds tests covering the fixed
cases plus the TVShow-grandparent and standalone-episode paths.
Add Wikidata IdTypeMapping entries so People items resolve cross-site:
P9650 -> IGDB_Company, P12836 -> DoubanPersonage (both were unmapped,
so claims on Wikidata entities never produced child People resources).

Add Spotify_Artist site class (open.spotify.com/artist/<id>, DEFAULT_MODEL
People) so the existing P1902 mapping can actually produce a scrapable
child resource instead of an orphan lookup_id. Uses Spotify /v1/artists.

Resolve the placeholder WIKI_PROPERTY_ID = "?" markers across site
classes: set real PIDs where one exists, "" where none does. Also fixes
two pre-existing wrong values picked up during the sweep:
  Goodreads class: P2968 (actually "QUDT unit ID") -> P2969
  MusicBrainzRelease class: P437 (actually "distribution format") -> P5813
My earlier insert of TestSpotifyArtist and TestIGDBCompany replaced only
TestGoodreadsAuthor.test_parse, leaving the original test_scrape block
below the new classes where Python re-parented it as TestIGDBCompany.
Move the Goodreads scrape back under its class. Caught by gemini-code-
assist in PR review.
Adds a button on a Person's page (when tmdb_person is set) that enqueues a
background job to fetch that person's full filmography from TMDB's
/person/{id}/combined_credits endpoint. Each Movie/TV entry is sent through
the existing enqueue_fetch pipeline, which dedupes by URL hash and skips
items already in the catalog. A per-person cache lock (10 min) prevents
repeat clicks while a batch is running.

No circular fetch: the task only enqueues work fetches; Movie.sync_credits_
from_metadata never enqueues back into person-works, so the graph terminates
at one level.
…l tasks

- Use cache.add() for the per-person fetch-works lock so the set-if-not-
  exists check is atomic, eliminating the race window where two concurrent
  requests could both pass the check and enqueue duplicate jobs.
- Pass user.pk (not the full User instance) into crawl-queue tasks and look
  the User up inside the task. Applies to fetch_works_for_person_task and
  also updates fetch_episodes_for_season_task to the same convention.
- Hoist the (previously inline) cache, User, enqueue_fetch, and
  tmdb_person_combined_credit_urls imports to module scope.
Prevents a single prolific author from dominating the recent/popular
posts list computed by the Discover cron.
Slice the queryset to limit * 10 before iterating so the database can
apply Top-N sort optimization; the headroom covers the per-author cap.
Search / shelf / catalog API:
- catalog/search/utils.py: replace Edition.works.first() with iteration of
  the prefetched .all() cache; .first() rebuilt the queryset and bypassed
  the prefetch, firing a per-edition ORDER BY pk LIMIT 1 (NEODB-SOCIAL-4BR).
- catalog/models/item.py: add Item.prefetch_edition_works() for ItemSchema
  callers that read parent_uuid unconditionally (shelf + catalog search APIs
  serialize every edition); keep prefetch_parent_items skipping Edition
  because the templates short-circuit edition.parent_item.
- journal/apis/shelf.py, catalog/apis.py: call prefetch_edition_works()
  alongside prefetch_parent_items (NEODB-SOCIAL-4HT).
- catalog/views/search.py, catalog/apis.py: batch prefetch credits with
  select_related(person) so search result cards don't fire a per-item
  itemcredit query (NEODB-SOCIAL-4JX, EGGPLANT-1AA).

Item detail pages:
- catalog/views/view.py: add _prefetch_reviews() paralleling
  _prefetch_comments; batches ratings and latest posts and attaches takahe
  identities for reviewers so /item/.../reviews stops firing per-review
  journal_rating / users_identity / journal_piecepost queries
  (NEODB-SOCIAL-4HM).
- catalog/views/view.py: in _prefetch_comments, select_related("parent") on
  the ShelfMember prefetch so mark.action_label no longer polymorphic-loads
  a Shelf per comment (NEODB-SOCIAL-4JW), and call the new takahe-identity
  batcher so comment.owner.display_name stops firing users_identity per
  comment (NEODB-SOCIAL-4HE).

Profile and collection pages:
- takahe/utils.py: add Takahe.prefetch_takahe_identities() to batch-populate
  APIdentity.takahe_identity.
- journal/views/profile.py: prefetch pinned_collections' owners with user
  chain, and pre-resolve identity.featured_collections so the sidebar's
  is_visible_to/get_stats loop stops firing users_apidentity/users_user/
  users_identity per featured collection (NEODB-SOCIAL-476).
- journal/models/collection.py: in get_members_by_page, set
  member.parent = self so collection_items.html does not polymorphic-load
  a Collection per rendered member (NEODB-SOCIAL-4CC).
- takahe/utils.py: cache Takahe.get_node_name_for_domain via django cache
  (10 min TTL) so ExternalResource.site_label stops firing one users_domain
  lookup per Fediverse-linked item on the collection page (EGGPLANT-17A).

Tests:
- tests/journal/test_n_plus_one.py: add regression tests asserting the
  shelf API makes no per-edition catalog_work queries, the reviews view
  makes no per-review rating or identity queries, and the collection view
  makes at most one journal_collection polymorphic lookup.
The catalog /search view calls Mark.attach_to_items(...) which batches
shelfmember/comment/rating/review/tag lookups and stashes a fully
populated Mark on each item as item.mark. _list_item.html resolves the
per-row mark through {% get_mark_for_item item as mark %}, which always
built a fresh Mark(user.identity, item) and ignored the prefetched copy,
so every access of mark.shelfmember / rating_grade / comment / review /
tags fired its own query (Sentry observed 60 repeats per render).

Reuse item.mark when it is attached and owned by the viewing identity;
fall back to constructing a fresh Mark when nothing was prefetched.
…schemas

The `_TVCreditResolverMixin` and `_PerformanceCreditResolverMixin` were
plain classes, so ninja's `ResolverMetaclass` never picked up their
`resolve_*` methods (it only inspects the schema class's own namespace
plus base classes that already expose `_ninja_resolvers`). Schema
serialization silently fell back to the legacy jsondata fields, and
`item.ap_object` validation raised `Input should be a valid list` when
those fields held corrupted scalar data.

Make both mixins inherit from `ninja.Schema` so the metaclass registers
the resolvers as intended, and add regression tests covering TVSeason,
TVShow, Performance, PerformanceProduction, and Movie.

Fixes NEODB-SOCIAL-4MQ
`CollectionForm.brief` defaulted to `required=True`, so saving a
collection with an empty description raised `BadRequest("Invalid
parameter")` even though `Collection.brief` is `TextField(blank=True,
default="")` and the rest of the save path handles an empty string.
Sentry flagged EGGPLANT-1B5 and EGGPLANT-1AS: each item card on
/collection/{uuid} fired its own catalog_itemcredit JOIN, and
_sidebar.html on /users/{user_name}/ ran CollectionMember and
ShelfMember GROUP BY queries once per featured collection.

Add Item.prefetch_credits() so Collection.get_members_by_page can
populate the credits prefetch cache that _credits_with_person already
checks. Add Collection.attach_stats_for_viewer() to precompute stats
for many collections in two queries, with get_stats short-circuiting
on the cache. Call both from the profile view.

Fixes EGGPLANT-1B5
Fixes EGGPLANT-1AS
Adopts review feedback to make the lookup defensive even though the
preceding loop initializes every ShelfType key.
…dits

Edition.pub_house and Edition.imprint are scalar `jsondata.CharField`s,
but `sync_credits_from_metadata` and `EditForm.canonicalize_credit_initials`
both treated every CREDIT_FIELD_MAPPING entry as a list. The form
re-populated the input with `str(['Foo'])` → "['Foo']", and saving back
persisted that literal string into the jsondata, which then survived
every subsequent edit.

- Track whether the original field is a string in both call sites and
  write back a single string for scalar fields.
- Add `resolve_pub_house`/`resolve_imprint` so the API reads the
  publisher/imprint from `credit_names_by_role(...)` (templates already
  do via `Edition.publisher_name`).

Existing items whose jsondata pub_house/imprint was already corrupted
into the literal "['Foo']" string still need a manual edit to clean —
the form now displays the raw stored value untouched, and a save with
brackets removed will round-trip correctly.
… lists

`Edition.pub_house` and `Edition.imprint` were scalar `jsondata.CharField`s
mapped through `CREDIT_FIELD_MAPPING`, which collided with both the form
canonicalization path and `sync_credits_from_metadata` (both assumed
list-typed values). Form round-trips therefore corrupted single-publisher
Editions into the literal string `"['Foo']"`.

Convert both to `jsondata.JSONField` lists (matching `author`/`translator`):

- Rename `pub_house` -> `publisher` on the model and in scrapers
  (`douban_book`, `goodreads`, `google_books`, `bangumi`, `bibliotek_dk`,
  `bookstw`, `openlibrary`, `worldcat`); each now writes
  `data["publisher"] = [name]`.
- Retype `imprint` to a list. `EditionInSchema.publisher: list[str]` and
  `imprint: list[str]` now resolve from credits. Keep the legacy
  `pub_house: str | None` schema field for API back-compat as a deprecated
  resolver returning `"/".join(credits)`.
- Add `Item.normalize_legacy_metadata(metadata)` hook called from
  `create_from_external_resource` and `merge_data_from_external_resource`.
  Edition's override translates `pub_house` (string), scalar `imprint`,
  and the `"['Foo']"` literal-corruption pattern via `ast.literal_eval`
  back into clean lists. Covers federated peers running older code and
  existing local rows.
- Async data migration `edition_publisher_imprint_to_list_20260428`
  (enqueued from `catalog/migrations/0023_*.py`) walks every Edition,
  normalizes metadata in bulk, drops the dead `pub_house` key, and then
  resyncs credits per touched item so fresh installs (where
  `populate_credits_20260412` ran before this migration) still get the
  rows.

Tests updated to use the list-shaped fields plus a new
`TestEditionLegacyMetadataCoercion` covering string, list, list-repr
corruption, scalar imprint, and the precedence/empty-key edge cases.
The previous refactor moved both `pub_house` and `imprint` onto the
credit system. Reconsidering: imprint is a single label per edition with
no person/organization semantics, so it does not belong in
`ItemCredit`. Revert the imprint half of that change while keeping the
publisher-as-list change intact.

- `Edition.imprint` reverts to `jsondata.CharField`; removed from
  `CREDIT_FIELD_MAPPING`. Schema field is `imprint: str | None` again
  (no `resolve_imprint`); template renders the scalar directly.
- `bookstw` and `douban_book` scrapers emit a scalar `imprint` again.
- `normalize_legacy_metadata` now collapses any list-shaped or
  ``"['Foo']"`` literal-corruption imprint metadata back into a scalar
  (joined with `/` if multiple), so federated peers and old backups
  carrying the intermediate list shape land cleanly.
- `PeopleRole.IMPRINT` removed; the `0022_alter_itemcredit_role`
  choices already exclude it, so no schema migration is needed.
- `populate_credits_extra_20260415` no longer maps Edition imprint into
  ItemCredit, so fresh installs never create those rows.
- `edition_normalize_publisher_imprint_20260428` is the one-off cleanup
  for instances that already ran the prior code: backfill the imprint
  scalar from any existing imprint credit rows, then delete the orphan
  `role="imprint"` ItemCredit rows. Function lives in
  `catalog/common/migrations.py`; no migration file is enqueued -- run
  manually post-deploy.

Tests updated to assert the scalar shape and the list/literal-repr
coercion.
Address review feedback on PR #1489:

- Replace OFFSET-based `Paginator` with keyset pagination
  (`pk__gt=last_pk`) so per-batch cost stays flat on large catalogs.
- Tighten the publisher-changed condition so rows where `publisher`
  is missing (`None`) and `pub_house` is also missing don't get marked
  dirty -- previously every such row was bulk-updated and queued for
  credit resync, the dominant case in production.
- Batch the post-update credit resync via `Edition.objects.filter(pk__in=chunk)`
  instead of one `Edition.objects.get(pk=pk)` round-trip per touched
  edition.
- common/views_manage.py: use getattr for is_superuser to handle
  AbstractBaseUser | AnonymousUser union from user_passes_test
- takahe/models.py: annotate User.REQUIRED_FIELDS and Post.objects as
  ClassVar to satisfy LSP against Django base classes
- common/models/jsondata.py: drop unused blanket type: ignore
Your Name and others added 28 commits May 24, 2026 10:09
…nking

In the sibling-dedup branch of fetch_linked_resources, update_content
only writes the new ExternalResource's metadata; the sibling Person's
own localized_name remains stale. _link_requester_credits reads
sibling.metadata.localized_name, so credits on the requester item whose
names only appear in the freshly-scraped data would be missed.

Merge new_res into the sibling before linking, mirroring the
get_item() path that calls merge_data_from_external_resource when an
existing item is matched. People has no CREDIT_FIELD_MAPPING, so the
embedded sync_credits_from_metadata call is a no-op.

Add a regression test that stubs TMDB_Person.scrape with multi-lingual
content (test settings restrict real scraping to English only) and
verifies a Chinese-named credit gets linked once the merge brings the
Chinese name into the sibling's localized_name.
Concurrent POSTs to /mark/{item_uuid} (double-submits, retries) both
pass the pre-check in append_item and race on the parent+item unique
constraint, leaking an IntegrityError to the user. Two fixes:

- List.append_item wraps the insert in transaction.atomic() and on
  IntegrityError re-fetches the existing member, returning it as
  (member, False) - append is already idempotent for the caller.
- TagManager.tag_item_for_owner uses get_or_create on Tag to close the
  same race on the Tag(owner, title) unique constraint that funnels
  concurrent requests onto the same parent.

Adds regression tests for both the race-recovery and idempotency paths.

Fixes NEODB-SOCIAL-3JG
Ensure manage.py shell sessions see current site-wide config and apply
it to settings, matching how RQ jobs and the request cycle bootstrap it.
Instruments auto-crosspost (Bluesky/Threads/Mastodon) and Mastodon
boost paths to emit crosspost.success / crosspost.failure counters
with platform and mode attributes, so dashboards can track delivery
to external social accounts and spot regressions.
Adds a small record_activity helper and instruments user-facing
entry points for post, article, mark, review, note and collection
in both API and web views. Importers are not instrumented.
Threads.post_single does not catch network errors, so non-RequestAborted
exceptions escaped sync_to_threads without emitting the failure metric.
Mirror sync_to_bluesky's catch-all to keep the counter reliable.
Audioless RSS items (chapter trailers, promo placeholders) cluttered
the episode list with nothing playable. Filter them out before insert,
and ship a `catalog prune-podcast-no-audio` action to hard-delete
existing audioless episodes that no user has journaled against.
Aligns both pyprojects ahead of a future single-pyproject merge so
remaining differences are just disjoint dep sets.

Dependency alignment:
- Python 3.14 across root + neodb-takahe (pyprojects, .python-version,
  Dockerfile, CI matrices in tests.yml/takahe.yml/check.yml, plus
  default_language_version in .pre-commit-config.yaml so hook envs
  can parse PEP 758 syntax).
- psycopg[binary]>=3.3.4 replaces psycopg2-binary in neodb.
- blurhash-rs>=1.1.0 replaces blurhash-python in both projects
  (cp314 wheels, ~250x faster; 2 call sites updated).
- pyld fork rebased onto upstream master @ 1889cab; rev pinned to
  a4834df on neodb-social/pyld:neodb-on-master-1889cab.
- Matched floors across both projects for cachetools (5.5.0),
  gunicorn (23.0.0), httpx (0.27.2), urlman (2.0.2),
  django (>=5.2,<5.3), django-storages (1.14.5).
- Takahe also gets cryptography 48.0.0 and pillow 12.2.0 floors
  for cp314 wheels.
- uv.lock regenerated with --upgrade in both projects.

Type fixes (ty 0.0.39 surfaced 10 diagnostics; all resolved):
- common/models/lang.py: widen _StrPromise-bearing function returns
  to dict[str, Any] / list[tuple[str, Any]].
- catalog/models/item.py: annotate _CATEGORY_LIST and skip None-category
  subclasses.
- catalog/common/sites.py: remove dead ResourceContent.dict() method
  (shadowed the builtin in class scope).
- journal/models/mark.py: cached_property notes returns concrete list.
- journal/models/common.py: widen Debris.create_from_piece arg to
  Content | ListMember.
- journal/models/utils.py: add isinstance(p, (Content, ListMember))
  guard before Debris.create_from_piece, matching the existing guard
  at line 72 and preventing a latent crash on ShelfLogEntry which
  lacks visibility/ap_object.

PEP 758 auto-format:
- 27 occurrences across 18 files: except (E1, E2): -> except E1, E2:
  (Python 3.14 makes parens optional; auto-applied by ruff).
neodb-takahe is now committed directly into the repo rather than tracked
as a git submodule, so the clone/update instructions and architecture
notes in development.md no longer need submodule steps.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
AbstractSite.query_str did `content.xpath(query)[0].strip()`, raising an
unhandled IndexError when the xpath matched nothing. The IMDB, Goodreads
and bibliotek.dk scrapers already guard the missing-__NEXT_DATA__ case
with `if not src: raise ParseError(...)`, but that guard was dead code
because query_str crashed before returning a falsy value.

Return "" on an empty match so those ParseError guards fire and are
handled (fetch_episodes_for_season skips the episode). Non-empty matches
are unchanged.

Fixes NEODB-SOCIAL-4H7

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
query_str now returns "" instead of raising IndexError on a missing
element. Bandcamp.scrape relied on catching that IndexError to reject
pages without a title/artist; without it, invalid pages would be saved
with empty title/artist metadata.

Check the parsed title/artist explicitly and raise ValueError when
either is empty, preserving the original rejection behavior. Add a
regression test covering a page missing title/artist.

Refs NEODB-SOCIAL-4H7

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Merge neodb-takahe's dependencies into the root pyproject.toml and drop its
separate pyproject/uv.lock, so both Django apps resolve from a single lockfile
and share one venv (/neodb-venv in Docker, .venv locally).

- root pyproject: add takahe deps (cryptography, pydantic, pillow, pyld,
  uvicorn, whitenoise, pywebpush, redis, rq, ...); consolidate django-storages
  extras to [boto3,google]; add pyld git source; fold takahe test deps into the
  dev group
- relax cryptography to >=46 (atproto caps it <47; takahe only uses stable
  hazmat APIs) so one version satisfies both apps
- move takahe pytest config to neodb-takahe/pytest.ini; delete
  neodb-takahe/pyproject.toml and neodb-takahe/uv.lock
- Dockerfile: build a single /neodb-venv from the unified project; takahe
  collectstatic and runtime use it
- compose.yml: TAKAHE_VENV -> /neodb-venv
- .github/workflows/takahe.yml: sync at the repo root, run takahe pytest via the
  shared project
- uv lock --upgrade (atproto 0.0.67, redis 8.0, django-redis 6.0, sentry-sdk
  2.61, idna 3.17, boto3 1.43.18, ...)

Verified in Docker with the single venv (Python 3.14.5):
neodb 1407 passed, takahe 355 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Several site scrapers could build a localized_title / localized_description
(or feed detect_language) with a None/empty value, which fails
LocalizedLabelSchema (lang and text must be non-null strings) during
item.ap_object validation and aborts the fetch -- the same class of bug as
the Qidian localized_description fix (NEODB-SOCIAL-7KK), found by auditing
all of catalog/sites.

Real crash paths (localized_title / localized_description, validated by the
ninja schema and feeding display_title/display_description):
- google_books: `language` is [] when a volume omits its language tag, so
  the localized_title `lang` was an empty list (not a str). Use a guaranteed
  string label_lang (declared language or detect_language(title)). Reachable
  on any volume without a language tag.
- discogs (Release + Master): title from `.get("title")` could be None ->
  raise ParseError instead of emitting text=None.
- igdb (Game): guard missing/None `name` with ParseError before use.
- bangumi: skip falsy titles when building localized_title (avoids
  detect_language(None) and text=None).
- mobygames: skip null alternateName entries.
- spotify (web + api): raise ParseError on missing/None name.
- worldcat: coerce None name/description before detect_language concat.

Cosmetic (localized_subtitle is not in EditionSchema and subtitle is
nullable, so these did not crash, but stored {text: None}):
- bookstw: subtitle is always None -> guard so localized_subtitle is [].
- douban_book: guard localized_subtitle like its sibling description.

Adds a google_books regression test (volume without a language tag); it
fails on the pre-fix code with `localized_title.0.lang` validation error
and passes with the fix.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Address review feedback on #1577:
- worldcat: a JSON-LD `inLanguage` list (common for multi-language books)
  made normalize_language() return None, so `lang` was None and failed
  LocalizedLabelSchema. Only normalize when in_language is a str, and fall
  back to detection otherwise.
- mobygames: guard that an `alternateName` entry is a str before
  detect_language() to avoid a TypeError on non-string JSON-LD values.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Qidian scraper unconditionally emitted
`[{"lang": "zh-cn", "text": brief}]`. When the description xpath
matched nothing, `brief` was None, producing a `{"text": None}` entry
that fails EditionSchema validation (None is not a valid string). On a
re-fetch the merge appended the bad entry after an existing description,
surfacing as `localized_description.1.text` in Sentry.

Guard the entry behind a truthiness check, matching the convention
already used by every other site scraper (goodreads, openlibrary, imdb,
douban_*, storygraph, bgg, google_books, etc.).

Adds a regression test with a fixture page that has a title but no
description; it fails on the pre-fix code with the exact validation
error and passes with the fix.

Fixes NEODB-SOCIAL-7KK

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
typesense 2.0 switched its HTTP transport from requests to httpx, so after
exhausting node retries it re-raises httpx exceptions (ConnectError,
TimeoutException, HTTPError, RequestError) that are not subclasses of
requests.RequestException or TypesenseClientError. With
connection_timeout_seconds=2, transient Typesense slowdowns raised these
straight up the stack (Sentry NEODB-SOCIAL-7KZ and similar), crashing search
views, federation post indexing and index-update workers.

Centralize the catchable error set in a TYPESENSE_ERRORS tuple
(RequestException, TypesenseClientError, httpx.HTTPError, JSONDecodeError) and
apply it to every runtime data operation (search, replace_docs, insert_docs,
delete_docs, patch_docs, get_doc). Also guard search() response parsing against
malformed responses, and harden get_doc() against transport errors while
preserving its ObjectNotFound not-found contract.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- search(): validate the multi_search response is a non-empty list whose
  first element is a dict before constructing SearchResult, so a malformed
  response (e.g. a string in "results") cannot raise AttributeError/TypeError.
- insert_docs(): return a consistent int document count (0 on empty input or
  transport error, count of successful inserts otherwise), mirroring
  replace_docs() instead of the previous mix of False/None/implicit-None.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…153)

Point the pyld source at neodb-social/pyld@e1de216 (branch
neodb-on-master-a7d380f), rebased onto upstream digitalbazaar/pyld
master@a7d380ff. The new commit makes the per-ResolvedContext processed
cache thread-safe, fixing "RuntimeError: OrderedDict mutated during
iteration" raised from concurrent canonicalise() in core/ld.py under
runstator's ThreadPoolExecutor. The earlier fork (a4834df / EGGPLANT-7K)
only protected the module-level caches, not the per-context one.

Fixes EGGPLANT-153

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a category dropdown to the user tag list and tag member pages so a
user's tags, and the items within a tag, can be filtered by item
category. The tag list shows category-scoped member counts and hides
tags that have no items in the selected category.

- TagManager.get_tags(category=...) annotates per-category counts and
  drops empty tags
- render_list filters tagmember querysets by category
- shared sidebar filter template for both pages; People and Collection
  are excluded (they cannot be tagged as catalog items)
- rename the tag list heading "All Tags" -> "Tag List"

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Use item_categories().get(category, []) (and skip classes missing from
item_content_types()) so a future ItemCategory without a corresponding
Item subclass yields an empty filter instead of raising KeyError.

All current ItemCategory values have Item subclasses, so behavior is
unchanged; this is defensive hardening from PR review.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Set both the project and patch coverage statuses to informational so
Codecov reports coverage on PRs but never posts a failing (red X) check.
Codecov has no lines-changed gate, so this is the practical lever to stop
small/low-coverage diffs from blocking PRs while still surfacing coverage
data in the PR comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Each catalog edit form now shows a genre list scoped to the item's
category instead of one flat ~95-item list. Movie and TV share a combined
screen list; Music, Game, Podcast, and Performance each get their own
(overlaps allowed).

Per-category lists live in common/models/genre.py (DEFAULT_GENRE_CATEGORIES,
grouped slug constants) and are overridable per category in the admin UI via
new SiteConfig.SystemOptions fields (genres_<category>) and a
"Catalog -> Genres by Category" settings page. GenreListField now takes a
category and builds its django-jsonform schema as a callable, so admin edits
take effect at render time without a restart. Out-of-list/custom values still
fall through the "Other" free-text option and normalize_genres unchanged.

No DB migration: genre fields are non-concrete (stored in the metadata JSON
column).
Address PR review: make genre_choices_for() robust to a manually
miscased category string (e.g. "Movie") by lowercasing the lookup key.
ItemCategory values are already lowercase, so existing call sites are
unaffected.
visible_categories caches the category list in the session, where
ItemCategory members are JSON-serialized to plain strings. On any
request after the first, the template received strings, so c.value and
c.label resolved to empty and the dropdown rendered blank options.

Switch to the membership-test pattern used by _sidebar_user_mark_list
(hardcoded option values + {% trans %} labels), which works whether
cats holds enum members or strings. People and Collection stay excluded.

Extend the page test to hit the list page on a repeat request (the
session-cached path) and assert real option values/labels render.
Pass collapse_profile=1 to the sidebar include on the tag list and tag
member pages so the profile section starts collapsed, keeping the
category filter and tag content above the fold.
Reorder the five per-category default genre lists (screen, music, game,
performance, podcast) alphabetically so the edit-item genre dropdown is
sorted out of the box. The sort is static in code rather than at render
time, so admin per-category overrides keep their own order/list, and the
free-text "Other" option is still appended last by GenreListField.schema.
redis-py 8.0.0 changed the default RESP protocol from 2 to 3
(DEFAULT_RESP_VERSION and the Connection protocol default), so the
client now sends `HELLO 3` on every connection. Production's Redis
endpoint does not implement HELLO and returns "unknown command HELLO",
breaking the Django cache for cached views such as webfinger.

Pin redis to >=7.4,<8 to restore the RESP2 default and stop the HELLO
handshake.

Fixes NEODB-SOCIAL-7M2
@alphatownsman
alphatownsman deleted the fix/mirror-canonical-repo branch June 4, 2026 23:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants