Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -235,6 +235,7 @@ The archive data itself is usually stored outside the repo or under a local `acc
| [`docs/agent-integration.md`](docs/agent-integration.md) | How AI agents consume the archive as a context layer |
| [`docs/agent-bundle-spec.md`](docs/agent-bundle-spec.md) | Agent bundle v1.0 specification |
| [`docs/provenance.md`](docs/provenance.md) | Manifest layer: provenance, integrity, and transform tracing |
| [`docs/schema-reference.md`](docs/schema-reference.md) | Consolidated `tweet.json`, agent-bundle, and manifest field reference |
| [`docs/cookbook-claude.md`](docs/cookbook-claude.md) | How Claude reads the archive for summarization/citation |
| [`docs/cookbook-hermes.md`](docs/cookbook-hermes.md) | How Hermes builds trend reports from the archive |
| [`docs/use-cases.md`](docs/use-cases.md) | Real-world scenarios and workflows |
Expand Down
150 changes: 150 additions & 0 deletions docs/schema-reference.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,150 @@
# Schema reference

This page is the consolidated field reference for the three versioned data
surfaces in CiteSeal. The agent-bundle and manifest tables mirror the JSON
Schemas in [`schemas/`](../schemas/). The `tweet.json` validator is intentionally
small and structural rather than a full JSON Schema; its required and
recommended top-level fields come from [`tools/scripts/tweet_schema.py`](../tools/scripts/tweet_schema.py).

`Required` means the field is required by the relevant schema. `Recommended`
is used for the `tweet.json` fields that the validator recommends but does not
reject when absent. Nested array rows describe the shape of each item.

## `tweet.json` v1

The validator requires `tweet_id`, `tweet_url`, `author_handle`, and
`datetime_utc`. It recommends the other top-level fields below. Media and
export entries are arrays of objects; the validator checks their `file` paths
against the item directory and preserves additional metadata fields.

### Minimum valid example

This self-contained example includes every required and recommended top-level
field. Empty arrays are valid when the captured item has no media, exports, or
replies.

```json
{
"tweet_id": "1234567890123456789",
"tweet_url": "https://x.com/example/status/1234567890123456789",
"author_handle": "example",
"datetime_utc": "2026-01-01T00:00:00Z",
"text": "Example post text.",
"media": [],
"exports": [],
"datetime_local": "2026-01-01T00:00:00Z",
"components": [],
"replies": []
}
```

| Field | Type | Required | Description |
|---|---|---:|---|
| `tweet_id` | string | Yes | Stable source-item identifier. |
| `tweet_url` | HTTP(S) URL string | Yes | Original X/Twitter post URL. |
| `author_handle` | string | Yes | Author handle, normally without a leading `@`. |
| `datetime_utc` | ISO 8601 string | Yes | Capture time in UTC. |
| `text` | string | Recommended | Captured post text. |
| `media` | array of objects | Recommended | Media declared for the item. |
| `media[].file` | string | Optional | Relative or bare filename resolved under the media directories. |
| `media[].type` | string | Optional | Media kind such as `image`, `video`, `audio`, or `raw`. |
| `media[].alt_text` | string | Optional | Human- or source-provided alternative text. |
| `media[].sha256` | string | Optional | SHA-256 hash when one has been recorded. |
| `media[].source_url` | HTTP(S) URL string | Optional | Original media URL when available. |
| `exports` | array of objects | Recommended | Derived files associated with the item. |
| `exports[].file` | string | Optional | Export filename under `exports/`. |
| `exports[].type` | string | Optional | Export kind such as `text` or `pdf`. |
| `datetime_local` | ISO 8601 string | Recommended | Capture time in the configured local timezone. |
| `components` | array | Recommended | Content components present in the item, such as `text` or `images`. |
| `replies` | array of objects | Recommended | Related replies used to build thread relationships. |
| `replies[].tweet_id` | string | Optional | Identifier of a related reply. |

## Agent bundle v1.0

Source: [`schemas/agent_bundle.schema.json`](../schemas/agent_bundle.schema.json).

| Field | Type | Required | Description |
|---|---|---:|---|
| `bundle_version` | string, constant `1.0` | Yes | Version of the agent-bundle specification. |
| `item_id` | string | Yes | Identifier of the archived item. |
| `source_platform` | string: `x`, `twitter`, or `web` | Yes | Platform where the item originated. |
| `source_url` | URI string | Yes | Original item URL. |
| `captured_at` | string | Yes | Capture timestamp copied from `tweet.json.datetime_utc`. |
| `author_handle` | string | Yes | Author handle without a leading `@`. |
| `text_excerpt` | string | Yes | Full text or a truncated excerpt ending in `...`. |
| `text_full` | string | No | Full untruncated text when it differs from the excerpt. |
| `media` | array of objects | No | Media assets attached to the item. |
| `media[].file` | string | Yes | Original filename from `tweet.json`. |
| `media[].type` | string: `image`, `video`, `audio`, or `raw` | Yes | Media kind. |
| `media[].path` | string | Yes | Relative path to the copied media in the bundle. |
| `media[].alt_text` | string | No | Media alternative text, if available. |
| `media[].sha256` | string | No | Media SHA-256 hash when computed. |
| `ocr_text` | string | No | OCR text extracted from media. |
| `article_md_path` | string | No | Relative path to an article Markdown export. |
| `assets` | array of objects | Yes | Every file included in the bundle output directory. |
| `assets[].path` | string | Yes | Relative asset path inside the bundle. |
| `assets[].kind` | string: `metadata`, `media`, `ocr`, `article`, `export`, `context`, or `manifest` | Yes | Asset classification. |
| `assets[].size_bytes` | integer | No | Asset size in bytes. |
| `citation_label` | string | No | Suggested human-readable citation label. |
| `trust_flags` | object | No | Data-quality and completeness flags. |
| `trust_flags.validated` | boolean | No | Whether the source `tweet.json` passed validation. |
| `trust_flags.has_media` | boolean | No | Whether the bundle contains media. |
| `trust_flags.has_ocr` | boolean | No | Whether OCR text is available. |
| `trust_flags.has_article` | boolean | No | Whether an article export is available. |
| `trust_flags.media_verified` | boolean | No | Whether declared media files were found on disk. |
| `provenance` | object | Yes | Capture and export provenance. |
| `provenance.exported_at` | string | Yes | Export timestamp. |
| `provenance.export_tool` | string | Yes | Tool and version that produced the bundle. |
| `provenance.source_dir` | string | No | Original source directory. |
| `provenance.schema_version` | string | No | Version of the source `tweet.json` schema. |
| `related_items` | array of objects | No | Related replies, quotes, retweets, or thread items. |
| `related_items[].item_id` | string | Yes | Related item identifier. |
| `related_items[].relation` | string: `reply`, `quote`, `retweet`, or `thread` | Yes | Relationship to the bundled item. |

## Manifest v1.0

Source: [`schemas/manifest.schema.json`](../schemas/manifest.schema.json).

| Field | Type | Required | Description |
|---|---|---:|---|
| `manifest_version` | string, constant `1.0` | Yes | Version of the manifest schema. |
| `item_id` | string | Yes | Identifier of the archived item. |
| `source_platform` | string: `x`, `twitter`, or `web` | Yes | Platform where the item originated. |
| `source_url` | URI string | Yes | Original item URL. |
| `captured_at` | string | Yes | Original capture timestamp. |
| `generated_at` | string | Yes | Timestamp when the manifest was generated. |
| `generator` | string | Yes | Tool that generated the manifest. |
| `author_handle` | string | No | Author handle without a leading `@`. |
| `citation_label` | string | No | Human-readable citation label. |
| `files` | array of `fileEntry` objects | Yes | Complete item-directory file inventory. |
| `files[].path` | string | Yes | Relative file path from the item directory root. |
| `files[].kind` | string: `metadata`, `media`, `media_raw`, `export`, `ocr`, `article`, `manifest`, or `other` | Yes | File classification. |
| `files[].size_bytes` | integer | Yes | File size in bytes. |
| `files[].sha256` | string | Yes | SHA-256 hash in hexadecimal. |
| `files[].derived_from` | array of strings | No | Source paths used to derive the file. |
| `files[].transform_id` | string | No | Transform that produced the file. |
| `transforms` | array of `transform` objects | Yes | Ordered processing history for the item. |
| `transforms[].id` | string | Yes | Unique transform identifier. |
| `transforms[].step` | string: `capture`, `transcode`, `ocr`, `article_md`, `article_pdf`, `fix`, `validate`, `export_agent`, or `manifest` | Yes | Processing step type. |
| `transforms[].tool` | string | Yes | Tool or script that performed the step. |
| `transforms[].started_at` | string | Yes | Transform start timestamp. |
| `transforms[].completed_at` | string | No | Transform completion timestamp. |
| `transforms[].status` | string: `success`, `warning`, `error`, or `skipped` | Yes | Transform outcome. |
| `transforms[].inputs` | array of strings | No | Input paths consumed by the step. |
| `transforms[].outputs` | array of strings | No | Output paths produced by the step. |
| `transforms[].notes` | string | No | Human-readable transform notes. |
| `components` | array of strings | No | Content components present in the item. |
| `trust_flags` | object | No | Archive-integrity and data-quality flags. |
| `trust_flags.has_metadata` | boolean | No | Whether `tweet.json` exists. |
| `trust_flags.has_media` | boolean | No | Whether media files exist. |
| `trust_flags.has_exports` | boolean | No | Whether export files exist. |
| `trust_flags.has_ocr` | boolean | No | Whether OCR files exist. |
| `trust_flags.has_article` | boolean | No | Whether article Markdown exists. |
| `trust_flags.media_verified` | boolean | No | Whether declared media files exist on disk. |
| `trust_flags.all_files_hashed` | boolean | No | Whether every file has a valid SHA-256 hash. |
| `summary` | object | No | Aggregate counts and sizes for the item directory. |
| `summary.total_files` | integer | No | Number of inventoried files. |
| `summary.total_bytes` | integer | No | Sum of inventoried file sizes. |
| `summary.media_count` | integer | No | Number of media files. |
| `summary.export_count` | integer | No | Number of export files. |
| `summary.transform_count` | integer | No | Number of transforms. |
Loading