Your data should belong to you.
Today your photos belong to a photo app, your messages to a messenger, your notes to a note app, your AI conversations to whoever runs the assistant. You are the tenant; the app is the landlord.
A pod inverts that. It is a data space you host, addressed over HTTP, holding structured linked data. Apps and agents come to your data instead of keeping copies of it, and you decide who may read or write what — and can change your mind without losing anything.
This repository is the reference implementation. The specification is its own repository, sempods-spec — a core every pod implements and optional modules on top, with hand-written OpenAPI descriptions of the HTTP surface.
The split is not bookkeeping. A contract that lives inside one implementation is a contract nobody
can tell apart from that implementation's habits; a second implementer reading docs/ here would
have had to guess which parts were obligations and which were Kotlin.
The lower stack is intentionally familiar. This implementation uses existing HTTP, RDF, SPARQL, OAuth/OIDC and MCP machinery where it can; the sempods-specific work is the contract that makes those pieces behave as one pod: context-scoped data, server-side grant resolution, sandboxed query surfaces and the same authority model for apps, websites and agents.
-
Resources are HTTP URIs.
https://example.org/alice/events/summer-partyis both the identifier and the address. Dereference it and you get RDF — JSON-LD by default, other serializations by content negotiation. -
Every statement lives in exactly one context. A context is a named graph, and it is the permission boundary. Not the resource, not the property — the context. One concept carries the whole access-control model.
-
Permissions are grants on contexts:
<context-iri>#read,#write,#manage. Apps and agents obtain them through OAuth 2.1 with PKCE; an app's identity is its origin, nameddid:web:<host>— nothing is fetched, and in production the authorization code goes nowhere but that origin. A grant is durable server-side policy — it never travels inside a token. -
SPARQL, with the sandbox enforced by the server. A query sees exactly the contexts the caller may read, and writes reach exactly the contexts the caller may write. Client-supplied dataset clauses are not trusted.
-
/_system/*is the control plane — contexts, grants, media, retrieval, the OAuth surface. It is not reachable through ordinary RDF writes.
One resource can hold public and private properties in different contexts at the same URI. An anonymous reader sees the public ones — automatic Linked Open Data — an authorized reader sees more. Same identifier, different depth, no duplication.
0.x. The surface moves. That is what the leading zero is for.
What runs in production today: pod servers hosting real tenants; event organizers publishing their programme as Linked Open Data that anyone can dereference and query without authentication; applications reading from several pods at once; an MCP endpoint per pod plus a hosted MCP service that fronts many pods, including pods run by other people. A second implementation of this contract runs inside another organisation's stack, built to its own architecture — its own tenancy, its own authorisation, its own search engine — against this documentation rather than against this code. That is the first evidence that the contract is implementable somewhere else, which is the claim this project actually needs to support.
Which of the two is right, while both exist. The specification is descriptive until it tags
0.1: it was extracted from this code, so where the two disagree today, this code is right. At
that tag it reverses, and a deviation here becomes the bug. gradle.properties names the version
this implements, and ./gradlew checkDocLinks fails if that claim drifts from the vendored index in
gradle/spec/.
What does not exist yet: a conformance suite — so nobody can prove an implementation conformant, including this one — and a one-command distribution.
Stable despite 0.x. A leading zero is a licence to move the API, not the data. These do not
move, because the first deployment that is not mine freezes them whatever the version number says
— changing one later costs a migration in somebody else's database, not a recompile:
- Package names.
org.sempods.*, already moved once, and not again. - Stored formats. Mongo database and collection names, the shape of the refresh-token rows, and the claim names in the tokens the services issue.
- Ontology IRIs. Every term under
https://schema.sempods.org/— see the vocabulary in sempods-spec, which also states the deprecation period they carry.
The one surface deliberately not on that list is the Maven coordinates. Snapshots are published, but a snapshot is mutable and expires; the coordinates freeze at the first release.
And one thing that is not stable, with no mechanism behind it. There is no schema migration
system. SempodsUpdater runs a hardcoded list on every boot; an update can declare itself
blocking and then finishes before the first request, which is the part that works. What is
missing is around it: no history, no "already applied" check, and a failure — blocking or not — is
logged while boot continues. The list holds one entry today.
sempods-server/docs/collections.md §"Schema changes" says
what that means for an upgrade.
Take a backup first, and read the startup log.
Who wrote it: one person, over years, with substantial AI assistance in the last of them. The security-relevant paths — the SPARQL sandbox, grant resolution, the OAuth flows — have had the most scrutiny, and independent review of them is the contribution I would value most. The project's subject is access control; it should be held to that standard rather than taken on trust.
You need Java 25 and Docker with the Compose plugin — the quick start's first command is
docker compose (v2), not the standalone docker-compose. Either binary works if you adjust the
line; podman-compose does too.
# 1. the only infrastructure a pod server needs
docker compose -f deployments/local/compose.yaml up -d
# 2. local configuration — one active line, which arms the development admin credential
cp deployments/local/env/local.example.env deployments/local/env/local.env
# 3. the pod server → http://localhost:8090
./gradlew :deployments:sempods:image:runCreate a pod and check it answers:
# the credential is the one step 2 armed. It is published with this source, so the server
# only accepts it when SEMPODS_DEV_ADMIN_FALLBACK asks; a deployment sets SEMPODS_ADMIN_CLIENTS
# instead, and without either every admin route answers 503.
#
# The owner is given as an email and stored as the WebID derived from it — sempods knows
# persons only as WebID URIs. → 201 {"pod":"demo","result":"created"}
curl -X PUT http://localhost:8090/_system/admin/pods/demo \
-H "Authorization: Bearer sc_development-admin-secret" \
-H "Content-Type: application/json" \
-d '{"ownerEmail":"alice@example.org"}'
# → 200 {"pod":"demo","exists":true}
curl http://localhost:8090/_system/admin/pods/demo \
-H "Authorization: Bearer sc_development-admin-secret"
# the pod's OAuth metadata — no authentication needed (RFC 9728)
curl http://localhost:8090/demo/.well-known/oauth-protected-resourceFrom here, docs/auth/oauth.md walks through registering an app and
obtaining a token, and the specification's CRUD chapter covers reading
and writing resources.
Configuration is documented where it is used; the variables that matter for a first run are
SEMPODS_HTTP_PORT, SEMPODS_PUBLIC_BASE_URL (the address the server is known by — pod IRIs
are minted from it), and MONGODB_URL. The natural-language layer has no off switch:
AI_PROVIDER defaults to ollama, so the AI routes are mounted in every deployment. Without a
token they answer 401, never 404 — which is how you can tell they are there — and with one they
answer 500 ai_provider_error until a provider is reachable. Everything else runs without one.
The libraries are on Maven Central as of 0.1.0. One version covers the whole repository, so pin
the platform and let the modules carry no version of their own:
dependencies {
implementation(platform("org.sempods:sempods-bom:0.1.0"))
implementation("org.sempods:sempods-client")
implementation("org.sempods:sempods-model")
}The modules are built, tested and released in lockstep, and a consumer holding sempods-client
0.2 against sempods-model 0.1 has a combination nothing ever ran — which is what the platform is
for, and why hand-versioning them is the one thing to avoid. What it carries are ordinary
constraints, so a different dependency asking for a newer sempods module can still pull that one
ahead of the rest. enforcedPlatform(...) in place of platform(...) makes them strict and forces
the platform's versions on the whole graph instead. That choice is left to you on purpose — made
here, it would propagate to everyone.
Published bytecode targets Java 21. Building this repository needs 25; depending on it does not.
The test fixtures — testFixtures("org.sempods:sempods-server") and the two sempods-commons
ones — resolve from Gradle, which reads the capability that carries them. Maven has no notion of
that capability, and the fixtures' own test libraries are deliberately kept out of the published
POM so that an ordinary consumer does not inherit them. The jar itself is published under the
test-fixtures classifier, so a Maven build can reach it — by naming that classifier and
supplying those dependencies itself.
Between releases main carries a -SNAPSHOT version, and merging to it republishes that version
to https://central.sonatype.com/repository/maven-snapshots/, which a build has to add explicitly.
Two things skip the publish: a version without -SNAPSHOT, so a main that is mid-release
publishes nothing, and a commit that is no longer the tip once its run reaches the gate, so a burst
of merges leaves only the newest — the snapshot follows main, not each commit on the way. A
snapshot is mutable, unvalidated and removed by Central after 90 days; pin a release instead unless
you specifically want to find out early that something changed.
No service calls another in-process: they meet over HTTP and environment variables, and each
starts, stops and scales without the others. What they do share is libraries — all three build on
sempods-auth-core, the pod server and the hosted MCP on sempods-mcp-core — so a token, a scope
and an MCP tool mean the same thing in each of them rather than nearly the same thing. Run one, two
or all three.
| Service | What it does | Needed when |
|---|---|---|
pod server (sempods-server) |
The pod itself: CRUD, SPARQL, contexts, grants, media, per-pod MCP | always |
identity (sempods-auth) |
WebID registry and OIDC bridge — gives people an identity a pod can grant to | you want person identities rather than only app credentials |
hosted MCP (sempods-mcp) |
One MCP connection fronting many pods, including pods run by others | you want an AI client to reach several pods at once |
Each ships as a container image — ghcr.io/haed/sempods, …/sempods-auth, …/sempods-mcp — under
latest, which a deployment pulls and which moves, plus the short commit for a build from a clean
checkout. That commit is always the OCI revision label, so a running container answers what it is
without anything being pulled:
docker inspect --format '{{index .Config.Labels "org.opencontainers.image.revision"}}' <container><sha>-dirty means uncommitted changes, unknown that the build could not identify its commit;
neither gets a tag, since neither names one commit. The commit tag names the source and not the
bytes — the base image floats, so rebuilding one commit republishes that tag with different content.
Pin a digest where exact content is what matters. The images are pushed by hand — each service's
jib task, no workflow — so that label is the only record of which commit reached the registry.
sempods-commons/ sempods-commons-json/ sempods-commons-mongo/
sempods-commons-okhttp/ sempods-commons-jaxrs/ sempods-commons-ktor/
framework-free shared base; take only what you need
sempods-model/ the contract as code — service interfaces, URI builder, ontologies
sempods-server/ the pod server: RDF4J store, contexts, OAuth, SPARQL, AI layer, MCP
sempods-media-s3/ the S3 binding of the media seam
sempods-auth/ identity service (Ktor)
sempods-auth-core/ the OAuth machinery all three services share — framework-free
sempods-mcp/ hosted MCP service (Ktor)
sempods-mcp-core/ the tool catalog and execution both MCP surfaces share
sempods-client/ HTTP client implementing the contract against a remote pod
sempods-control-plane-client/
HTTP client for the host-level admin surface (pod hosting)
deployments/ the server as a process, and the local stack
docs/ how this implementation works, and why. The contract is sempods-spec
The server is a reference implementation, not one particular hosting. Behaviours a
deployment may need to replace — the RDF store, the find engine, resource expansion, the AI
provider, admin authority — are interfaces with a deployment-selected binding rather than
forks. Which ones exist, which do not yet, and what each costs is documented in
docs/concepts/modularity.md.
| sempods-spec | the contract — contexts, grants, auth, CRUD, SPARQL, find, and the three modules. Start there to implement a pod |
docs/ |
everything about this implementation, indexed as Vision, maintained IST documentation and occasional Proposal |
docs/vision.md |
the model and why it is shaped this way |
docs/auth/ |
what this implementation does around the OAuth contract: rate limits, timeouts, provisioning, the error page |
docs/mcp/ |
this implementation's MCP surfaces: the tool reference, the challenge store, and how real clients behave |
docs/concepts/graph-retrieval.md |
graph retrieval — find, then traverse |
docs/media.md |
the media storage seam: which backends exist, how a deployment picks one |
docs/concepts/modularity.md |
what a deployment may replace |
docs/concepts/ |
retained architecture and design material awaiting classification |
| GitHub issues | public goals, planned iterations, decisions and progress |
Maintained documentation describes current code; issues own public plans, with occasional explicitly proposed design documents. Existing roadmap and SOLL content follows the bounded transition. The documentation strategy defines ownership and completion for human and AI contributors alike.
Small changes are welcome. Follow the
issue-planning rules for the work record and
completion evidence, including automated updates and private security fixes. Agree larger changes
before implementation.
Contributions run under the Developer Certificate of Origin — git commit -s — and there is
deliberately no CLA: everyone, maintainer included, works under the same licence. See
CONTRIBUTING.md, which also lists the handful of properties that will not
change.
Response times vary. This is not yet anyone's full-time job.
Vulnerability reports go to hello@sempods.org, never into a public issue. See
SECURITY.md for scope, expectations, and the design decisions that look like
vulnerabilities but are not.
Code is licensed Apache 2.0 (LICENSE). Documentation, the specification and
the vocabulary are CC BY 4.0.
The Apache licence grants no rights to the name (§6), so what you may call your own work is set
out separately in TRADEMARKS.md — deliberately permissive: build it, run it
commercially, embed it in a closed product, fork it. The name is regulated only where it would
suggest that this project produced or endorsed something it did not.
Vocabulary terms and their stability guarantees: sempods-spec vocabulary/.
Questions, ideas, or interest in building on this: hello@sempods.org