Recast is a small REST API for creating AI characters (personas) and chatting with them. Each character has a name, a bio, and a personality. When you send a message, the API wraps your text in a system prompt that tells the underlying Large Language Model (LLM) to "stay in character," then returns the reply.
The same character behaves consistently no matter which LLM provider you use — swap between OpenAI, Anthropic (Claude), Google (Gemini), or Groq with a single request parameter, and "Detective Rao" still talks like Detective Rao.
Built with FastAPI + SQLAlchemy + PostgreSQL, packaged with uv and Docker Compose. Zero manual setup:
docker compose upand you're live.
- What you get
- Quick start (one command)
- Setting your API keys (
.env) - Try it in the browser (
/docs) - Switching providers
- API reference
- How it works (architecture)
- New concepts, explained
- Project layout
- Demo
- Known limitations & roadmap
- ✅
docker compose upstarts the API + database with no manual steps — no installing Python, nopip install, no running migrations by hand. - ✅ Create a character, chat with it, and fetch history — all from the
interactive
/docspage in your browser. - ✅ Switch LLM providers with one parameter (
provider), and the character's persona stays consistent. - ✅ Secrets stay out of git —
.envis.gitignored, and this README shows you exactly how to set your keys.
Prerequisites: Docker Desktop (which includes Docker Compose). That's the only thing you need installed.
# 1. Clone the repo
git clone <your-repo-url> recast
cd recast
# 2. Create your .env from the template and add at least one API key (see below)
cp .env.example .env # then edit .env
# 3. Start everything
docker compose upThat's it. Two containers come up:
| Service | What it is | Address |
|---|---|---|
api |
The Recast FastAPI app | http://localhost:8000 |
db |
PostgreSQL 16 database | localhost:5432 |
The database schema (tables for characters, conversations, and messages) is created automatically on startup, so there is nothing to migrate or seed.
Open http://localhost:8000/docs to start using the API.
Stopping: press
Ctrl+C, thendocker compose down. Your data persists in a Docker volume (pgdata); usedocker compose down -vto wipe it.
Recast talks to external LLM providers, and each provider needs an API key.
Keys are secrets — they must never be committed to git. We keep them in a
file called .env (already listed in .gitignore), and the app
loads them at startup.
Create your .env by copying the template:
cp .env.example .envThen open .env and fill in the key(s) for the provider(s) you want to use.
You only need a key for the provider you actually call — Groq is the default
because it has a generous free tier.
# .env — never commit this file
# Database (the default already matches docker-compose; you can leave it as-is)
DATABASE_URL=postgresql://postgres:postgres@db:5432/persona
# LLM provider keys — fill in the one(s) you'll use, leave the rest blank
GROQ_API_KEY=gsk_... # default provider, free tier — https://console.groq.com
OPENAI_API_KEY=sk-... # https://platform.openai.com/api-keys
ANTHROPIC_API_KEY=sk-ant-... # https://console.anthropic.com
GEMINI_API_KEY=... # https://aistudio.google.com/apikeyWhere do these names come from? The app reads environment variables in
app/config.pyusing pydantic-settings. The variable names are matched case-insensitively, soGROQ_API_KEYin.envmaps to thegroq_api_keysetting in code.
A ready-to-copy template is provided as .env.example so you
never have to guess the variable names.
FastAPI ships with Swagger UI — an interactive, auto-generated web page that
lists every endpoint and lets you call them with a form. No curl, no Postman.
Open http://localhost:8000/docs and follow these three steps:
Expand POST /characters → Try it out → paste a body → Execute:
{
"name": "Detective Rao",
"bio": "A sharp, world-weary homicide detective in 1980s Mumbai.",
"personality": "Dry wit, speaks in clipped sentences, suspicious of everyone, quietly kind."
}The response includes the new character's id (e.g. 1) — remember it.
Expand POST /chat → Try it out:
{
"user_id": "alice",
"character_id": 1,
"message": "Detective, where were you on the night of the murder?",
"provider": "groq"
}You'll get a reply in character:
{
"reply": "Working. Always working. The city doesn't sleep, so neither do I. Why do you ask?",
"provider": "groq"
}Expand GET /conversations/{user_id}/{character_id}, enter alice and 1,
and Execute to retrieve the stored message history for that user + character.
Prefer raw HTTP? There's also a ReDoc view and the machine-readable spec at
/openapi.json.
The provider field on POST /chat chooses which LLM answers. Everything else —
the character, the system prompt, the conversation — stays identical, so the
persona is consistent across providers. Just change one word:
provider |
Backend | Model used |
|---|---|---|
groq (default) |
Groq | llama-3.3-70b-versatile |
openai |
OpenAI | gpt-4o-mini |
claude |
Anthropic | claude-sonnet-4-20250514 |
gemini |
gemini-2.0-flash |
The routing logic lives in app/llm.py. Each provider has a
slightly different API shape (for example, Anthropic takes the system prompt as a
separate argument rather than as a message), and chat_llm() normalizes those
differences behind one function.
| Method & path | Description |
|---|---|
POST /characters |
Create a character. Body: name, bio, personality, voice_id. |
GET /characters |
List all characters. |
GET /characters/{cid} |
Fetch one character by id. |
DELETE /characters/{cid} |
Delete a character by id. |
POST /chat |
Send a message to a character and get an in-character reply. |
GET /conversations/{user_id}/{character_id} |
Fetch the message history for a user + character. |
# Create
curl -X POST localhost:8000/characters \
-H "Content-Type: application/json" \
-d '{"name":"Detective Rao","bio":"1980s Mumbai detective.","personality":"Dry, clipped, suspicious."}'
# Chat
curl -X POST localhost:8000/chat \
-H "Content-Type: application/json" \
-d '{"user_id":"alice","character_id":1,"message":"Who did it?","provider":"groq"}'
# History
curl localhost:8000/conversations/alice/1 ┌────────────────────────────────────────┐
HTTP request │ Recast API │
───────────────▶ /chat │ (FastAPI, app/main.py) │
│ │
│ 1. look up Character in PostgreSQL │
│ 2. build a system prompt from its │
│ name + personality │
│ 3. call chat_llm(provider, …) │──▶ OpenAI / Groq /
│ 4. (persist messages — see roadmap) │ Claude / Gemini
│ 5. return the reply │◀── reply
└───────────────┬────────────────────────┘
│ SQLAlchemy ORM
▼
┌────────────────────┐
│ PostgreSQL (db) │
│ characters │
│ conversations │
│ messages │
└────────────────────┘
Request lifecycle for POST /chat:
- FastAPI validates the JSON body against the
ChatInmodel. crud.get_character()loads the character from the database (404 if missing).system_prompt()turns the character's name + personality into instructions that force the model to stay in character.chat_llm()dispatches to the chosen provider and returns the reply text.- The reply is returned as JSON.
The data model (app/models.py) has three tables:
- Character — the persona (
name,bio,personality,voice_id). - Conversation — one thread between a
user_idand acharacter_id. - Message — a single turn (
role="user"or"assistant", pluscontent), belonging to a conversation.
If some of the tools here are unfamiliar, here's a plain-language tour.
A modern Python web framework for building APIs. You write a function, decorate it
with @app.post("/chat"), and FastAPI handles routing, request parsing, validation,
and generates the interactive /docs page for free.
Pydantic defines the shape of data. CharacterIn says "a valid character-creation
request has a name (string) and optional bio/personality." FastAPI uses these
to validate incoming JSON and to document the API. Invalid requests are rejected
automatically with a helpful 422 error — you never write that validation by hand.
A companion to Pydantic that loads configuration from environment variables and the
.env file into a typed Settings object. This is how secrets get from .env into
the app without being hard-coded.
An ORM (Object-Relational Mapper) lets you work with database rows as Python
objects instead of writing raw SQL. models.Character is a Python class that maps
to the characters table; db.add(obj) inserts a row. Base.metadata.create_all()
creates all the tables on startup.
Notice db: Session = Depends(get_db) in the endpoints. FastAPI calls get_db()
for each request, hands the resulting database session to your function, and closes
it afterward — automatic per-request resource management.
LLMs accept a system prompt: high-priority instructions that set the model's
behavior before the user's message. Recast builds one per character
(see system_prompt() in app/main.py) telling the model who it is
and to never reveal it's an AI. This is what makes a persona consistent across
different providers.
Each LLM vendor has its own SDK and message format. chat_llm() is a thin
adapter that presents one uniform interface and translates to each provider's
specifics internally — so the rest of the app doesn't care which model answers.
A fast Python package manager and resolver (a modern replacement for pip +
virtualenv). uv sync --frozen installs the exact dependency versions pinned in
uv.lock, giving reproducible builds. You don't run it directly — the Docker image
does.
Docker packages the app and its dependencies into a portable image (built from
the Dockerfile). Docker Compose (docker-compose.yml)
runs multiple containers together — here, the API and the database — and wires
them up (networking, environment, startup order via depends_on) so a single
docker compose up brings the whole system online.
Recast/
├── app/
│ ├── main.py # FastAPI app + all HTTP endpoints
│ ├── llm.py # chat_llm(): routes to OpenAI / Groq / Claude / Gemini
│ ├── models.py # SQLAlchemy tables: Character, Conversation, Message
│ ├── schemas.py # Pydantic request/response models
│ ├── crud.py # database read/write helpers
│ ├── db.py # engine, session factory, get_db() dependency
│ └── config.py # settings loaded from .env
├── Dockerfile # builds the API image with uv
├── docker-compose.yml# api + db services
├── pyproject.toml # project metadata & dependencies
├── uv.lock # pinned, reproducible dependency versions
└── .env # your secrets (gitignored — create from .env.example)
🎬 Watch a chat with "Detective Rao": <add your Loom/GIF link here>
This is an early build. A couple of things are scaffolded but not yet fully wired:
- Conversation persistence in
/chatis stubbed. Inapp/main.py, the/chathandler currently sends an empty history to the model and does not save the user message or the reply to the database (see the# ...placeholder comments). As a result,GET /conversations/...returns an empty list until this is implemented. The database tables and the read endpoint already exist — what's left is to create/load theConversation, pull the last ~10Messagerows as history, and persist both turns. - No authentication yet —
user_idis taken at face value from the request. - Local/dev defaults — the Postgres credentials in
docker-compose.ymlare for local development only; use real secrets in any deployed environment.
Contributions and fixes welcome.