From c902e280199c67f68d9817b5d3ac7fe7dd9ee8f8 Mon Sep 17 00:00:00 2001 From: "Aryan Singh K." <70511529+aryansk@users.noreply.github.com> Date: Sun, 9 Aug 2026 13:48:07 +0530 Subject: [PATCH 1/2] docs: add consolidated schema reference --- README.md | 1 + docs/schema-reference.md | 129 +++++++++++++++++++++++++++++++++++++++ 2 files changed, 130 insertions(+) create mode 100644 docs/schema-reference.md diff --git a/README.md b/README.md index 129940d..4d24ece 100644 --- a/README.md +++ b/README.md @@ -235,6 +235,7 @@ The archive data itself is usually stored outside the repo or under a local `acc | [`docs/agent-integration.md`](docs/agent-integration.md) | How AI agents consume the archive as a context layer | | [`docs/agent-bundle-spec.md`](docs/agent-bundle-spec.md) | Agent bundle v1.0 specification | | [`docs/provenance.md`](docs/provenance.md) | Manifest layer: provenance, integrity, and transform tracing | +| [`docs/schema-reference.md`](docs/schema-reference.md) | Consolidated `tweet.json`, agent-bundle, and manifest field reference | | [`docs/cookbook-claude.md`](docs/cookbook-claude.md) | How Claude reads the archive for summarization/citation | | [`docs/cookbook-hermes.md`](docs/cookbook-hermes.md) | How Hermes builds trend reports from the archive | | [`docs/use-cases.md`](docs/use-cases.md) | Real-world scenarios and workflows | diff --git a/docs/schema-reference.md b/docs/schema-reference.md new file mode 100644 index 0000000..4825667 --- /dev/null +++ b/docs/schema-reference.md @@ -0,0 +1,129 @@ +# Schema reference + +This page is the consolidated field reference for the three versioned data +surfaces in CiteSeal. The agent-bundle and manifest tables mirror the JSON +Schemas in [`schemas/`](../schemas/). The `tweet.json` validator is intentionally +small and structural rather than a full JSON Schema; its required and +recommended top-level fields come from [`tools/scripts/tweet_schema.py`](../tools/scripts/tweet_schema.py). + +`Required` means the field is required by the relevant schema. `Recommended` +is used for the `tweet.json` fields that the validator recommends but does not +reject when absent. Nested array rows describe the shape of each item. + +## `tweet.json` v1 + +The validator requires `tweet_id`, `tweet_url`, `author_handle`, and +`datetime_utc`. It recommends the other top-level fields below. Media and +export entries are arrays of objects; the validator checks their `file` paths +against the item directory and preserves additional metadata fields. + +| Field | Type | Required | Description | +|---|---|---:|---| +| `tweet_id` | string | Yes | Stable source-item identifier. | +| `tweet_url` | HTTP(S) URL string | Yes | Original X/Twitter post URL. | +| `author_handle` | string | Yes | Author handle, normally without a leading `@`. | +| `datetime_utc` | ISO 8601 string | Yes | Capture time in UTC. | +| `text` | string | Recommended | Captured post text. | +| `media` | array of objects | Recommended | Media declared for the item. | +| `media[].file` | string | Optional | Relative or bare filename resolved under the media directories. | +| `media[].type` | string | Optional | Media kind such as `image`, `video`, `audio`, or `raw`. | +| `media[].alt_text` | string | Optional | Human- or source-provided alternative text. | +| `media[].sha256` | string | Optional | SHA-256 hash when one has been recorded. | +| `media[].source_url` | HTTP(S) URL string | Optional | Original media URL when available. | +| `exports` | array of objects | Recommended | Derived files associated with the item. | +| `exports[].file` | string | Optional | Export filename under `exports/`. | +| `exports[].type` | string | Optional | Export kind such as `text` or `pdf`. | +| `datetime_local` | ISO 8601 string | Recommended | Capture time in the configured local timezone. | +| `components` | array | Recommended | Content components present in the item, such as `text` or `images`. | +| `replies` | array of objects | Recommended | Related replies used to build thread relationships. | +| `replies[].tweet_id` | string | Optional | Identifier of a related reply. | + +## Agent bundle v1.0 + +Source: [`schemas/agent_bundle.schema.json`](../schemas/agent_bundle.schema.json). + +| Field | Type | Required | Description | +|---|---|---:|---| +| `bundle_version` | string, constant `1.0` | Yes | Version of the agent-bundle specification. | +| `item_id` | string | Yes | Identifier of the archived item. | +| `source_platform` | string: `x`, `twitter`, or `web` | Yes | Platform where the item originated. | +| `source_url` | URI string | Yes | Original item URL. | +| `captured_at` | string | Yes | Capture timestamp copied from `tweet.json.datetime_utc`. | +| `author_handle` | string | Yes | Author handle without a leading `@`. | +| `text_excerpt` | string | Yes | Full text or a truncated excerpt ending in `...`. | +| `text_full` | string | No | Full untruncated text when it differs from the excerpt. | +| `media` | array of objects | No | Media assets attached to the item. | +| `media[].file` | string | Yes | Original filename from `tweet.json`. | +| `media[].type` | string: `image`, `video`, `audio`, or `raw` | Yes | Media kind. | +| `media[].path` | string | Yes | Relative path to the copied media in the bundle. | +| `media[].alt_text` | string | No | Media alternative text, if available. | +| `media[].sha256` | string | No | Media SHA-256 hash when computed. | +| `ocr_text` | string | No | OCR text extracted from media. | +| `article_md_path` | string | No | Relative path to an article Markdown export. | +| `assets` | array of objects | Yes | Every file included in the bundle output directory. | +| `assets[].path` | string | Yes | Relative asset path inside the bundle. | +| `assets[].kind` | string: `metadata`, `media`, `ocr`, `article`, `export`, `context`, or `manifest` | Yes | Asset classification. | +| `assets[].size_bytes` | integer | No | Asset size in bytes. | +| `citation_label` | string | No | Suggested human-readable citation label. | +| `trust_flags` | object | No | Data-quality and completeness flags. | +| `trust_flags.validated` | boolean | No | Whether the source `tweet.json` passed validation. | +| `trust_flags.has_media` | boolean | No | Whether the bundle contains media. | +| `trust_flags.has_ocr` | boolean | No | Whether OCR text is available. | +| `trust_flags.has_article` | boolean | No | Whether an article export is available. | +| `trust_flags.media_verified` | boolean | No | Whether declared media files were found on disk. | +| `provenance` | object | Yes | Capture and export provenance. | +| `provenance.exported_at` | string | Yes | Export timestamp. | +| `provenance.export_tool` | string | Yes | Tool and version that produced the bundle. | +| `provenance.source_dir` | string | No | Original source directory. | +| `provenance.schema_version` | string | No | Version of the source `tweet.json` schema. | +| `related_items` | array of objects | No | Related replies, quotes, retweets, or thread items. | +| `related_items[].item_id` | string | Yes | Related item identifier. | +| `related_items[].relation` | string: `reply`, `quote`, `retweet`, or `thread` | Yes | Relationship to the bundled item. | + +## Manifest v1.0 + +Source: [`schemas/manifest.schema.json`](../schemas/manifest.schema.json). + +| Field | Type | Required | Description | +|---|---|---:|---| +| `manifest_version` | string, constant `1.0` | Yes | Version of the manifest schema. | +| `item_id` | string | Yes | Identifier of the archived item. | +| `source_platform` | string: `x`, `twitter`, or `web` | Yes | Platform where the item originated. | +| `source_url` | URI string | Yes | Original item URL. | +| `captured_at` | string | Yes | Original capture timestamp. | +| `generated_at` | string | Yes | Timestamp when the manifest was generated. | +| `generator` | string | Yes | Tool that generated the manifest. | +| `author_handle` | string | No | Author handle without a leading `@`. | +| `citation_label` | string | No | Human-readable citation label. | +| `files` | array of `fileEntry` objects | Yes | Complete item-directory file inventory. | +| `files[].path` | string | Yes | Relative file path from the item directory root. | +| `files[].kind` | string: `metadata`, `media`, `media_raw`, `export`, `ocr`, `article`, `manifest`, or `other` | Yes | File classification. | +| `files[].size_bytes` | integer | Yes | File size in bytes. | +| `files[].sha256` | string | Yes | SHA-256 hash in hexadecimal. | +| `files[].derived_from` | array of strings | No | Source paths used to derive the file. | +| `files[].transform_id` | string | No | Transform that produced the file. | +| `transforms` | array of `transform` objects | Yes | Ordered processing history for the item. | +| `transforms[].id` | string | Yes | Unique transform identifier. | +| `transforms[].step` | string: `capture`, `transcode`, `ocr`, `article_md`, `article_pdf`, `fix`, `validate`, `export_agent`, or `manifest` | Yes | Processing step type. | +| `transforms[].tool` | string | Yes | Tool or script that performed the step. | +| `transforms[].started_at` | string | Yes | Transform start timestamp. | +| `transforms[].completed_at` | string | No | Transform completion timestamp. | +| `transforms[].status` | string: `success`, `warning`, `error`, or `skipped` | Yes | Transform outcome. | +| `transforms[].inputs` | array of strings | No | Input paths consumed by the step. | +| `transforms[].outputs` | array of strings | No | Output paths produced by the step. | +| `transforms[].notes` | string | No | Human-readable transform notes. | +| `components` | array of strings | No | Content components present in the item. | +| `trust_flags` | object | No | Archive-integrity and data-quality flags. | +| `trust_flags.has_metadata` | boolean | No | Whether `tweet.json` exists. | +| `trust_flags.has_media` | boolean | No | Whether media files exist. | +| `trust_flags.has_exports` | boolean | No | Whether export files exist. | +| `trust_flags.has_ocr` | boolean | No | Whether OCR files exist. | +| `trust_flags.has_article` | boolean | No | Whether article Markdown exists. | +| `trust_flags.media_verified` | boolean | No | Whether declared media files exist on disk. | +| `trust_flags.all_files_hashed` | boolean | No | Whether every file has a valid SHA-256 hash. | +| `summary` | object | No | Aggregate counts and sizes for the item directory. | +| `summary.total_files` | integer | No | Number of inventoried files. | +| `summary.total_bytes` | integer | No | Sum of inventoried file sizes. | +| `summary.media_count` | integer | No | Number of media files. | +| `summary.export_count` | integer | No | Number of export files. | +| `summary.transform_count` | integer | No | Number of transforms. | From c5a94f10f447f841cd90a2fc47d5b856427082cf Mon Sep 17 00:00:00 2001 From: "Aryan Singh K." <70511529+aryansk@users.noreply.github.com> Date: Sun, 9 Aug 2026 20:04:44 +0530 Subject: [PATCH 2/2] docs: add minimum tweet schema example --- docs/schema-reference.md | 21 +++++++++++++++++++++ 1 file changed, 21 insertions(+) diff --git a/docs/schema-reference.md b/docs/schema-reference.md index 4825667..a763678 100644 --- a/docs/schema-reference.md +++ b/docs/schema-reference.md @@ -17,6 +17,27 @@ The validator requires `tweet_id`, `tweet_url`, `author_handle`, and export entries are arrays of objects; the validator checks their `file` paths against the item directory and preserves additional metadata fields. +### Minimum valid example + +This self-contained example includes every required and recommended top-level +field. Empty arrays are valid when the captured item has no media, exports, or +replies. + +```json +{ + "tweet_id": "1234567890123456789", + "tweet_url": "https://x.com/example/status/1234567890123456789", + "author_handle": "example", + "datetime_utc": "2026-01-01T00:00:00Z", + "text": "Example post text.", + "media": [], + "exports": [], + "datetime_local": "2026-01-01T00:00:00Z", + "components": [], + "replies": [] +} +``` + | Field | Type | Required | Description | |---|---|---:|---| | `tweet_id` | string | Yes | Stable source-item identifier. |