os2ai/Feedback#18
Enable token tracking in the openwebui admin analytics UI.
Tracking can be turned on manually in the UI for each model and litellm will deliver the amount of tokens used.
Default enabling of usage/token tracking forwarded to openwebui from litellm seems to only be possible by enabling the litellm general setting always-include-streaming-usage: true
https://docs.litellm.ai/docs/completion/usage#proxy-always-include-streaming-usage
Alternatives that have been tried but did not work:
The issue with setting always-include-streaming-usage: true might be the following:
- If OS2ai chooses to switch from openwebui to some other UI might be that the UI does not support this and crashes. We need to ensure that the UI works for that.
- There might be circumstances where providing usage numbers introduces some overhead which means longer responses for end-users.
An alternative solution to feeding the usage stats into openwebui might be introducing the litellm UI as something the end-user admins should use instead. This would decrease the technical complexity, but will introduce additional GRC work for the end-users.
The https://demo.os2ai.dk is running with always-include-streaming-usage: true as a test.
os2ai/Feedback#18
Enable token tracking in the openwebui admin analytics UI.
Tracking can be turned on manually in the UI for each model and litellm will deliver the amount of tokens used.
Default enabling of usage/token tracking forwarded to openwebui from litellm seems to only be possible by enabling the litellm general setting
always-include-streaming-usage: truehttps://docs.litellm.ai/docs/completion/usage#proxy-always-include-streaming-usage
Alternatives that have been tried but did not work:
include_usageinDEFAULT_MODEL_PARAMSin openwebui. https://docs.openwebui.com/reference/env-configuration/#default_model_params os2ai/demo-gitops@36bc5b9include_usageon the individual models in litellm. os2ai/demo-gitops@756583cThe issue with setting
always-include-streaming-usage: truemight be the following:An alternative solution to feeding the usage stats into openwebui might be introducing the litellm UI as something the end-user admins should use instead. This would decrease the technical complexity, but will introduce additional GRC work for the end-users.
The https://demo.os2ai.dk is running with
always-include-streaming-usage: trueas a test.