Skip to content

fix(api): honour encoding_format=base64 in /v1/embeddings - #5646

Merged
qinxuye merged 1 commit into
xorbitsai:mainfrom
MohammadHijjawi97:fix/embeddings-base64-encoding-format
Oct 7, 2026
Merged

qinxuye merged 1 commit into
xorbitsai:mainfrom
MohammadHijjawi97:fix/embeddings-base64-encoding-format

Conversation

@MohammadHijjawi97

Copy link
Copy Markdown
Contributor

What

POST /v1/embeddings accepted encoding_format but dropped it and always returned float arrays.

  • A client that sets encoding_format="base64" gets a list where the OpenAI API returns a base64 string, so code that decodes the string breaks.
  • The OpenAI Python and Node SDKs ask for base64 by default and decode it client-side, so every SDK call to Xinference pays for the float payload. For a 1024-dim vector that is ~22 KB per embedding instead of ~5.5 KB.

Fix

When encoding_format == "base64", each dense vector is encoded the way OpenAI does it: little-endian float32 bytes, base64-encoded. Sparse embeddings (token -> weight dicts) have no base64 form and stay as they are. Default and "float" responses are unchanged.

Before / after (openai SDK 3.26.0 against the handler)

r = await client.embeddings.create(model="m", input="hi", encoding_format="base64")
r.data[0].embedding
# before: [0.5, -1.0, 2.0]
# after:  'AAAAPwAAgL8AAABA'   (decodes to [0.5, -1.0, 2.0])

The SDK default call (create(model=..., input=...)) still returns [0.5, -1.0, 2.0]. It now gets base64 on the wire and decodes it.

Tests

  • New xinference/api/tests/test_embedding_encoding_format.py:
    • base64 output round-trips to the original float32 vectors, and encoding_format is not forwarded to the model.
    • Omitted and "float" keep lists.
    • Sparse embeddings are unchanged.
    • The base64 test fails without the fix.
  • pytest xinference/api/tests: all pass, plus pre-commit (black, ruff, isort, mypy, codespell) on the changed files.

The embeddings endpoint dropped encoding_format and always returned
float arrays. Clients that explicitly ask for base64 then receive a list
where the OpenAI contract promises a base64 string, and the OpenAI
Python and Node SDKs, which request base64 by default, always pay for
the larger float payload (about 4x for a 1024-dim vector).

Encode dense vectors as base64 little-endian float32 when base64 is
requested, matching OpenAI. Sparse embeddings have no base64 form and
are returned unchanged.
@XprobeBot XprobeBot added the bug Something isn't working label Oct 7, 2026
@XprobeBot XprobeBot added this to the v3.x milestone Oct 7, 2026

@qinxuye qinxuye left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Verified dense base64 encoding, unchanged float/default responses and sparse handling. All 14 focused tests passed, and an additional OpenAI SDK check passed for default, float and explicit base64 requests. The Metal CI failure is in F5-TTS model startup (actor connection reset), outside the changed embedding response path.

@qinxuye
qinxuye merged commit 9cd2c87 into xorbitsai:main Oct 7, 2026
14 of 15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants