Skip to content

Commit de1e79d

Browse files
authored
Refactor project (#45)
* feat: add additional script for en-only embeddings * feat: switch from flake8 to ruff * refactor: embedding deployment scripts * chore: code formatting * feat: pin public URL labels for embeddings endpoints
1 parent 0d0843f commit de1e79d

22 files changed

Lines changed: 1083 additions & 1027 deletions

‎.env.example‎

Lines changed: 13 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -1,18 +1,17 @@
1-
PHI_URL=....
2-
QWEN_URL=...
1+
# ── LLM endpoints (OpenAI-compatible vLLM servers on Modal) ──
2+
PHI_URL=https://<account>--phi-4-mini-instruct-qa-vllm-serve.modal.run/v1
3+
QWEN_URL=https://<account>--qwen-0-6b-qa-vllm-serve.modal.run/v1
4+
API_KEY=your-llm-api-key
35

4-
EMBEDS_URL=...
6+
# ── Embedding endpoint ───────────────────────────────────────
7+
EMBEDS_URL=https://<account>--intfloat-multilingual-e5-large-instruct-embeddings-embed.modal.run
8+
EMBEDS_API_KEY=your-embedding-api-key
9+
10+
# ── Defaults pre-selected in the UI ──────────────────────────
511
DEFAULT_MODEL=microsoft/Phi-4-mini-instruct
612
DEFAULT_EMBEDDING=intfloat/multilingual-e5-large-instruct-modal
713

8-
API_KEY=...
9-
EMBEDS_API_KEY=...
10-
11-
GROBID_URL=...
12-
GROBID_QUANTITIES_URL=...
13-
14-
15-
QWEN_URL=...
16-
GROBID_MATERIALS_URL=...
17-
API_KEY=...
18-
EMBEDS_API_KEY=...
14+
# ── GROBID services ──────────────────────────────────────────
15+
GROBID_URL=https://your-grobid-url
16+
GROBID_QUANTITIES_URL=https://your-grobid-quantities-url/ # optional (measurements NER)
17+
GROBID_MATERIALS_URL=https://your-grobid-superconductors-url/ # optional (materials NER)

‎.github/workflows/ci-build.yml‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -31,14 +31,14 @@ jobs:
3131
- name: Install dependencies
3232
run: |
3333
python -m pip install --upgrade pip
34-
pip install --upgrade flake8 pytest pycodestyle pytest-cov huggingface_hub
34+
pip install --upgrade ruff pytest pytest-cov huggingface_hub
3535
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
36-
- name: Lint with flake8
36+
- name: Lint with ruff
3737
run: |
3838
# stop the build if there are Python syntax errors or undefined names
39-
flake8 . --count --select=E9,F63,F7,F82 --show-source --statistics
40-
# exit-zero treats all errors as warnings. The GitHub editor is 127 chars wide
41-
flake8 . --count --exit-zero --max-complexity=10 --max-line-length=127 --statistics
39+
ruff check --select=E9,F63,F7,F82 --output-format=full .
40+
# non-blocking: report all remaining style issues without failing the build
41+
ruff check --exit-zero --statistics .
4242
- name: Test with pytest
4343
run: |
4444
pytest

‎.github/workflows/ci-release.yml‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -25,14 +25,14 @@ jobs:
2525
- name: Install dependencies
2626
run: |
2727
python -m pip install --upgrade pip
28-
pip install --upgrade flake8 pytest pycodestyle
28+
pip install --upgrade ruff pytest
2929
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
30-
- name: Lint with flake8
30+
- name: Lint with ruff
3131
run: |
3232
# stop the build if there are Python syntax errors or undefined names
33-
flake8 . --count --select=E9,F63,F7,F82 --show-source --statistics
34-
# exit-zero treats all errors as warnings. The GitHub editor is 127 chars wide
35-
flake8 . --count --exit-zero --max-complexity=10 --max-line-length=127 --statistics
33+
ruff check --select=E9,F63,F7,F82 --output-format=full .
34+
# non-blocking: report all remaining style issues without failing the build
35+
ruff check --exit-zero --statistics .
3636
# - name: Test with pytest
3737
# run: |
3838
# pytest

‎README.md‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -46,6 +46,8 @@ Additionally, this frontend provides the visualisation of named entities on LLM
4646

4747
**For full technical documentation** of the `document-qa-engine` library **[`docs/README.md`](docs/README.md)**.
4848

49+
**To deploy the LLM and embedding endpoints** on Modal.com, see **[`document_qa/deployment/README.md`](document_qa/deployment/README.md)**.
50+
4951
### Embedding selection
5052
In the latest version, there is the possibility to select both embedding functions and LLMs. There are some limitations, OpenAI embeddings cannot be used with open source models, and vice-versa.
5153

@@ -83,7 +85,7 @@ For more information, see the [details](https://docs.trychroma.com/troubleshooti
8385
Please read carefully:
8486

8587
- Avoid uploading sensitive data. We temporarily store text from the uploaded PDF documents only for processing your request, and we disclaim any responsibility for subsequent use or handling of the submitted data by third-party LLMs.
86-
- Mistral and Zephyr are FREE to use and do not require any API, but as we leverage the free API entrypoint, there is no guarantee that all requests will go through. Use at your own risk.
88+
- The public demo serves open models (Phi-4-mini-instruct, Qwen3) self-hosted on [Modal.com](https://www.modal.com) under a limited monthly compute budget, so there is no guarantee that all requests will go through. Use at your own risk.
8789
- We do not assume responsibility for how the data is utilized by the LLM end-points API.
8890

8991
## Development notes

‎docs/README.md‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -97,6 +97,13 @@ GROBID_MATERIALS_URL=https://your-grobid-superconductors-url/
9797
| `GROBID_QUANTITIES_URL` | URL to a grobid-quantities server (for measurement NER) |
9898
| `GROBID_MATERIALS_URL` | URL to a grobid-superconductors server (for materials NER) |
9999

100+
### Deploying the model endpoints
101+
102+
The `PHI_URL`, `QWEN_URL`, and `EMBEDS_URL` endpoints above are served by the Modal apps
103+
in [`../document_qa/deployment/`](../document_qa/deployment/README.md). That README covers
104+
the required secrets, deploy commands, and how each printed `*.modal.run` URL maps back to
105+
these variables.
106+
100107
---
101108

102109
## Quick Start — Streamlit App

‎document_qa/custom_embeddings.py‎

Lines changed: 13 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -47,18 +47,13 @@ def embed(self, text: List[str]) -> List[List[float]]:
4747
# Newlines degrade embedding quality for most models
4848
cleaned_text = [t.replace("\n", " ") for t in text]
4949

50-
payload = {'text': "\n".join(cleaned_text)}
50+
payload = {"text": "\n".join(cleaned_text)}
5151

5252
headers = {}
5353
if self.api_key:
54-
headers = {'x-api-key': self.api_key}
55-
56-
response = requests.post(
57-
self.url,
58-
data=payload,
59-
files=[],
60-
headers=headers
61-
)
54+
headers = {"x-api-key": self.api_key}
55+
56+
response = requests.post(self.url, data=payload, files=[], headers=headers)
6257
response.raise_for_status()
6358

6459
# print(response.text)
@@ -92,12 +87,15 @@ def get_model_name(self) -> str:
9287

9388

9489
if __name__ == "__main__":
90+
# Smoke test against a deployed Modal embedding endpoint. The endpoint requires
91+
# the x-api-key header, so set EMBEDS_URL and EMBEDS_API_KEY in the environment
92+
# (see document_qa/deployment/README.md).
93+
import os
94+
9595
embeds = ModalEmbeddings(
96-
url="https://lfoppiano--intfloat-multilingual-e5-large-instruct-embed-5da184.modal.run/",
97-
model_name="intfloat/multilingual-e5-large-instruct"
96+
url=os.environ["EMBEDS_URL"],
97+
model_name="intfloat/multilingual-e5-large-instruct",
98+
api_key=os.environ.get("EMBEDS_API_KEY"),
9899
)
99100

100-
print(embeds.embed(
101-
["We are surrounded by stupid kids",
102-
"We are interested in the future of AI"]
103-
))
101+
print(embeds.embed(["We are surrounded by stupid kids", "We are interested in the future of AI"]))

‎document_qa/deployment/README.md‎

Lines changed: 105 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,105 @@
1+
# Modal deployment scripts
2+
3+
This folder contains the [Modal](https://modal.com) apps that serve the LLM and
4+
embedding endpoints used by document-qa. Each script is an independent Modal app:
5+
deploy the ones you need, then point the matching `.env` variables at the URLs
6+
Modal prints.
7+
8+
| Script | Modal app | Serves | Maps to `.env` |
9+
|--------|-----------|--------|----------------|
10+
| `modal_inference_phi.py` | `phi-4-mini-instruct-qa-vllm` | `microsoft/Phi-4-mini-instruct` (vLLM, OpenAI-compatible) | `PHI_URL` |
11+
| `modal_inference_qwen.py` | `qwen-0.6b-qa-vllm` | `Qwen/Qwen3-0.6B` (vLLM, reasoning) | `QWEN_URL` |
12+
| `modal_embeddings_multilang.py` | `intfloat-multilingual-e5-large-instruct-embeddings` | `intfloat/multilingual-e5-large-instruct` | `EMBEDS_URL` |
13+
| `modal_embeddings_en.py` | `intfloat-e5-large-v2-embeddings` | `intfloat/e5-large-v2` (English-only) | `EMBEDS_URL` |
14+
15+
> Both embedding scripts define a tiny global `EmbeddingModel` class that delegates
16+
> to the shared helpers in `_embeddings_app.py` (`cls_kwargs`, `load_embedding_model`,
17+
> `run_embed`). The shared module holds the container image and the embedding logic;
18+
> the model is loaded **once per container** via `@modal.enter()`. To add another
19+
> embedding model, copy one wrapper and change `MODEL_NAME` / `MODEL_REVISION` / the
20+
> app name.
21+
22+
## Prerequisites
23+
24+
```bash
25+
pip install modal
26+
modal token new # one-time browser auth
27+
```
28+
29+
## Secrets
30+
31+
The scripts read an `API_KEY` from a Modal [Secret](https://modal.com/docs/guide/secrets).
32+
Create the two secrets once (the value is the bearer token clients must send):
33+
34+
```bash
35+
# Used by the inference scripts (phi, qwen)
36+
modal secret create document-qa-api-key API_KEY=<your-llm-token>
37+
38+
# Used by the embedding scripts
39+
modal secret create document-qa-embedding-key API_KEY=<your-embedding-token>
40+
```
41+
42+
| Secret | Used by | Provides |
43+
|--------|---------|----------|
44+
| `document-qa-api-key` | `modal_inference_phi.py`, `modal_inference_qwen.py` | `API_KEY` for the vLLM `--api-key` flag |
45+
| `document-qa-embedding-key` | `modal_embeddings_*.py` | `API_KEY` checked against the `x-api-key` header |
46+
47+
## Deploy
48+
49+
```bash
50+
modal deploy document_qa/deployment/modal_inference_phi.py
51+
modal deploy document_qa/deployment/modal_inference_qwen.py
52+
modal deploy document_qa/deployment/modal_embeddings_multilang.py
53+
# modal deploy document_qa/deployment/modal_embeddings_en.py # optional English-only
54+
```
55+
56+
Each deploy prints a public `https://<...>.modal.run` URL. Copy it into `.env`:
57+
58+
```env
59+
PHI_URL=https://<account>--phi-4-mini-instruct-qa-vllm-serve.modal.run/v1
60+
QWEN_URL=https://<account>--qwen-0-6b-qa-vllm-serve.modal.run/v1
61+
EMBEDS_URL=https://<account>--embeddings-multilang.modal.run # English-only: --embeddings-en
62+
API_KEY=<your-llm-token> # matches document-qa-api-key
63+
EMBEDS_API_KEY=<your-embedding-token> # matches document-qa-embedding-key
64+
```
65+
66+
> **Inference endpoints** are OpenAI-compatible vLLM servers, so their URLs end in
67+
> `/v1`. **Embedding endpoints** are a custom form endpoint (see below), so their
68+
> URL has no `/v1` suffix.
69+
70+
## Endpoint contracts
71+
72+
### Inference (vLLM)
73+
74+
Standard OpenAI Chat Completions API at `<PHI_URL|QWEN_URL>`, authenticated with the
75+
`Authorization: Bearer <API_KEY>` header. Used by `langchain_openai.ChatOpenAI` in
76+
`streamlit_app.py`.
77+
78+
### Embeddings
79+
80+
A custom `POST` endpoint consumed by
81+
[`ModalEmbeddings`](../custom_embeddings.py):
82+
83+
- **Auth**: `x-api-key: <EMBEDS_API_KEY>` header.
84+
- **Body**: form field `text` with newline-separated strings.
85+
- **Response**: JSON list of L2-normalised vectors, one per input line.
86+
87+
Smoke test:
88+
89+
```bash
90+
curl -X POST "$EMBEDS_URL" \
91+
-H "x-api-key: $EMBEDS_API_KEY" \
92+
-F $'text=first sentence\nsecond sentence'
93+
```
94+
95+
## Tuning
96+
97+
These knobs live near the top of each script (or in `_embeddings_app.py`):
98+
99+
| Setting | Where | Notes |
100+
|---------|-------|-------|
101+
| `gpu` | `@app.function` / `@app.cls` | `A10G` is cheaper; `L40S` is faster. Embeddings default to `L40S`, inference to `A10G`. |
102+
| `scaledown_window` | decorator | Idle time before a replica is stopped (cost vs. cold starts). |
103+
| `max_inputs` | `@modal.concurrent` | Concurrent requests per replica — tune to GPU memory. |
104+
| `LABEL` | `modal_embeddings_*.py` | Pins the public URL (`--<label>.modal.run`). Without it Modal truncates the long auto-name and appends a random hash. |
105+
| `FAST_BOOT` | `modal_inference_phi.py` | `--enforce-eager` for faster cold starts vs. peak throughput. |
Lines changed: 116 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,116 @@
1+
"""Shared building blocks for the Modal embedding endpoints.
2+
3+
``modal_embeddings_en.py`` and ``modal_embeddings_multilang.py`` each define a tiny
4+
``EmbeddingModel`` class at module scope (Modal requires globally-defined classes
5+
with stacked ``@app.cls`` / ``@modal.concurrent`` decorators) that delegates to the
6+
helpers here. All the heavy lifting — the container image, model loading, pooling,
7+
and the embedding request handler — lives in this module so it is written once.
8+
9+
The endpoint contract (consumed by ``document_qa.custom_embeddings.ModalEmbeddings``):
10+
11+
- **Method**: ``POST``
12+
- **Auth**: ``x-api-key`` header, compared against the ``API_KEY`` secret.
13+
- **Body**: form field ``text`` containing newline-separated strings.
14+
- **Response**: JSON list of L2-normalised embedding vectors, one per input line.
15+
"""
16+
17+
import os
18+
19+
import modal
20+
import torch
21+
import torch.nn.functional as F
22+
from fastapi import HTTPException, Request
23+
from torch import Tensor
24+
25+
MINUTES = 60 # seconds
26+
N_GPU = 1
27+
28+
# Shared container image for every embedding model.
29+
image = (
30+
modal.Image.debian_slim(python_version="3.11")
31+
.pip_install(
32+
"transformers",
33+
"huggingface_hub[hf_transfer]==0.26.2",
34+
"flashinfer-python==0.2.0.post2", # pinning, very unstable
35+
"fastapi[standard]",
36+
extra_index_url="https://flashinfer.ai/whl/cu124/torch2.5",
37+
)
38+
.env({"HF_HUB_ENABLE_HF_TRANSFER": "1"}) # faster model transfers
39+
# Modal 1.0 no longer auto-mounts imported local modules; the wrapper scripts
40+
# import this module by name, so it must be added explicitly. Kept last so it
41+
# doesn't invalidate the (expensive) pip layer above on every code edit.
42+
.add_local_python_source("_embeddings_app")
43+
)
44+
45+
hf_cache_vol = modal.Volume.from_name("huggingface-cache", create_if_missing=True)
46+
vllm_cache_vol = modal.Volume.from_name("vllm-cache", create_if_missing=True)
47+
48+
49+
def cls_kwargs() -> dict:
50+
"""Common ``@app.cls`` configuration shared by every embedding endpoint."""
51+
return dict(
52+
image=image,
53+
gpu=f"L40S:{N_GPU}",
54+
# how long should we stay up with no requests?
55+
scaledown_window=3 * MINUTES,
56+
volumes={
57+
"/root/.cache/huggingface": hf_cache_vol,
58+
"/root/.cache/vllm": vllm_cache_vol,
59+
},
60+
secrets=[modal.Secret.from_name("document-qa-embedding-key")],
61+
)
62+
63+
64+
def average_pool(last_hidden_states: Tensor, attention_mask: Tensor) -> Tensor:
65+
"""Mean-pool token embeddings, ignoring padding positions."""
66+
last_hidden = last_hidden_states.masked_fill(~attention_mask[..., None].bool(), 0.0)
67+
return last_hidden.sum(dim=1) / attention_mask.sum(dim=1)[..., None]
68+
69+
70+
def load_embedding_model(model_name: str, model_revision: str):
71+
"""Load a tokenizer + model onto the best available device, once per container.
72+
73+
Returns:
74+
tuple: ``(tokenizer, model, device)`` with ``model`` already in eval mode.
75+
"""
76+
# transformers is only available inside the Modal image, so import lazily.
77+
from transformers import AutoModel, AutoTokenizer
78+
79+
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
80+
print(f"Loading {model_name} on {device}...")
81+
tokenizer = AutoTokenizer.from_pretrained(model_name, revision=model_revision)
82+
model = AutoModel.from_pretrained(model_name, revision=model_revision).to(device)
83+
model.eval()
84+
print("Model loaded successfully.")
85+
return tokenizer, model, device
86+
87+
88+
def run_embed(tokenizer, model, device, request: Request, text: str):
89+
"""Authenticate, embed newline-separated ``text``, and return normalised vectors."""
90+
api_key = request.headers.get("x-api-key")
91+
if api_key != os.environ["API_KEY"]:
92+
raise HTTPException(status_code=401, detail="Unauthorized")
93+
94+
texts = [t for t in text.split("\n") if t.strip()]
95+
if not texts:
96+
return []
97+
98+
print(f"Start embedding {len(texts)} texts")
99+
try:
100+
with torch.no_grad():
101+
batch_dict = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
102+
batch_dict = {k: v.to(device) for k, v in batch_dict.items()}
103+
104+
outputs = model(**batch_dict)
105+
embeddings = average_pool(outputs.last_hidden_state, batch_dict["attention_mask"])
106+
embeddings = F.normalize(embeddings, p=2, dim=1)
107+
embeddings = embeddings.cpu().numpy().tolist()
108+
109+
print("Finished embedding texts.")
110+
return embeddings
111+
112+
except RuntimeError as e:
113+
print(f"Error during embedding: {str(e)}")
114+
if "CUDA out of memory" in str(e):
115+
print("CUDA OOM. Try reducing batch size or using a smaller model.")
116+
raise

0 commit comments

Comments
 (0)