Repository navigation
Conversation
Add `tabpfn.telemetry`, whose `log_usage` decorator turns each outermost call to an estimator's fit, predict and embedding methods into a usage event in the format of the telemetry API. Calls a logged method makes internally, such as the holdout fits of `tuning_config` or the training and validation fits of fine-tuning, are not logged on their own. A batched call is one event with the number of datasets it scored. Events record the sizes of the data and the estimator's settings, never the data itself, and only settings and values the API accepts. They go nowhere until a sink is installed with `set_sink`, which only happens for accounts that have opted in; delivery to the API follows separately. Wrap the entry points of TabPFNClassifier, TabPFNRegressor and the fine-tuned estimators with it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add `tabpfn.telemetry.collector`, modelled on the queue of PostHog's Python client: usage events go on a bounded in-memory queue, and a background thread sends them in batches of up to 100, once 10 are waiting or every 5 seconds. A batch that cannot be sent is saved under ~/.cache/tabpfn/telemetry and sent once the API answers again, from this process or a later one; events keep their id, so PostHog deduplicates a batch sent twice. The first logged call starts collecting if there is an API key. The background thread then asks `GET /account/telemetry` and sends or saves nothing unless the account opted in; if it did not, the thread ends and nothing more is logged. A 403 from `POST /telemetry` stops collecting the same way. No call waits on the network, and exit waits at most 2 seconds. Fix `check_telemetry_enabled` to call the route gapi defines, to return None when it cannot tell, and never to raise. Log the GPU each call ran on. A forked process logs nothing: on macOS its first network request can crash it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The warning that a column looks like free text names the user's `estimator.fit(X, y)` call by counting frames up to it. `log_usage` wraps `fit` in one more frame, so the warning named the wrapper in `tabpfn/telemetry/decorator.py` instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Saved batches were named after `time.time_ns()`, whose clock on Windows ticks only every 15ms, so batches saved within one tick sorted by their random suffix, and the oldest were not always the ones deleted at the limit. Number the batches a process saves, after the time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
brendan-priorlabs
left a comment
There was a problem hiding this comment.
Thank you, @safaricd. Left a couple comments for you. Requesting changes just to be cautious
| test_shapes = {_shape(X) for X in arguments["X_test_list"]} | ||
| fields: dict[str, Any] = { | ||
| # An empty list fails the call, and the API only accepts a positive count. | ||
| "num_datasets": len(arguments["X_train_list"]) or None, |
There was a problem hiding this comment.
Claude thinks there's a bug here with "batched" requests. Seems serious. Implies the need for and e2e with the server in the loop.
- Batched predictions are always rejected by the server, and take other
events down with them
- events.py:120 puts a num_datasets field on batched events (from
predict_proba_batched and predict_batched).- The server's event models don't define that field, and its base model is
extra="forbid" (libs/prior/models/pydantic.py:13). So the server rejects the
request with a 422.- The server accepts or rejects a batch as a whole
(routers/telemetry/telemetry.py). On the client, _send (collector.py:540)
treats a 4xx response as "done" and throws the batch away.- Result: every batched call silently discards up to 99 good events that were
sent with it.- The client tests assert num_datasets exists (test_telemetry.py:497), so they
never test against the server's schema.- Fix: add num_datasets to PredictCalled on the server, or drop it on the
client. Also add a test that checks client events against the server's
models
| SHUTDOWN_TIMEOUT = 2.0 | ||
| """Seconds that sending the queued events may delay the exit of the process.""" | ||
|
|
||
| SAVED_BATCHES_ROOT = CACHE_DIR / "telemetry" |
There was a problem hiding this comment.
How about we don't write anything to disk? We can use an exit handler to write everything at exit and set the batch size to be small enough that we don't lose too much.
There was a problem hiding this comment.
Dropped - events not dumped to disk any longer but rather only using the in-memory queue with a lower flush period.
| if not enabled: | ||
| # The saved batches are kept if the API could not say, to be sent | ||
| # by a process that hears that the account opted in. | ||
| self._disable(keep_saved=enabled is None) |
There was a problem hiding this comment.
We should err on the side of caution. If the API can't say, we don't record anything.
|
|
||
|
|
||
| @lru_cache(maxsize=1) | ||
| def check_telemetry_enabled(token: str, api_url: str) -> bool | None: |
There was a problem hiding this comment.
I have a question here regarding opt-in/out, etc. Will follow up directly.
Rename `tabpfn.telemetry` to `tabpfn.analytics`, and its wording with it. The API routes keep their names, `/telemetry` and `/account/telemetry`. Drop the batches saved to disk. A batch the API does not take is kept in memory and sent again after `RETRY_DELAY`, while later events wait in the bounded queue; events not sent when the process ends are dropped. A stop ends the wait to send again, so that the exit is not delayed. Record nothing when the API cannot say whether usage analytics is enabled for the account, as when it says it is not. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Changes
log_usagelogs each outermost call to an estimator's fit, predict and embedding methods as one usage event of data sizes and settings, never the data.TabPFNClassifier,TabPFNRegressorand the fine-tuned estimators are decorated; calls they make internally (tuning holdouts, fine-tuning fits,forward) are not logged again.predict_proba_batched,predict_batched) is logged as one event withnum_datasets.gpu_type), looked up once per device.tabpfn/analytics/collector.py, modelled on PostHog's client, queues events in memory and sends them from a background thread in batches of up to 100, once 10 are waiting or every 5 s.GET /account/telemetrybefore sending anything.POST /telemetryreturns 403, collecting stops and the queued events are dropped.check_telemetry_enablednow calls/account/telemetry, returnsTrue/False/Nonelikecheck_license_accepted, and never raises.stacklevelsteps past thelog_usagewrapper, so it still names the user'sfitcall.tests/conftest.pyturns logging off for the test suite.Motivation
Usage analytics for licensed accounts whose contract covers it, sent to gapi's
POST /telemetry. It is on only if an API key exists (TABPFN_TOKEN,~/.cache/tabpfn/auth_tokenor~/.tabpfn/token) andGET /account/telemetryreturns{"enabled": true}, which Prior Labs sets per account with the customer's consent. No environment variable turns it on or off; the API is the only switch. For everyone else nothing is sent, and no thread outlives the check. No call waits on the network, and exit waits at most 2 s. Requires gapi'snum_datasets/ optionalnum_rowsschema change to be deployed before release.How this was tested
tests/test_analytics.py: decorator and events (40 tests)forwardis logged once.torch.deviceobjects.other.model_path="auto"logs the default model version.output_type, quantiles and the regression task.TabPFNClassifier.predict.gpu_typeis the CUDA device's name, looked up once per device.gpu_typeis None on CPU andmpson Apple GPUs.gpu_typeis None when CUDA cannot start.gpu_typeis None for a device name gapi would reject.TabPFNClassifierlogs one event per entry-point call.TabPFNClassifierwithtuning_configlogs one fit.TabPFNRegressorlogs one event per entry-point call.tests/test_analytics_collector.py: delivery against a local fake gapi (25 tests)_postreturns the response status (204, 401, 403, 422, 503)._postreturns None when the connection is refused or the URL is malformed._postreturns None when the API answers too slowly.stop()sends queued events without waiting for the flush interval.stop()returns quietly when the thread never started.stop()while waiting to send a batch again returns at once.stop()while a request hangs returns within its timeout.start()andstop()deliver a decorated estimator's events end to end.tests/test_browser_auth.py::TestCheckTelemetryEnabled(5 tests){"enabled": true}counts as enabled; other bodies are False and invalid JSON is None./account/telemetry, with or without a trailing slash on the base URL.Schema, regressions and static checks (6 checks)
test__fit_with_text_column__transform_text_off__warns_at_call_sitepins the free-text warning to the user'sfitcall through the wrapper.pre-commit(ruff, ruff format, mypy) passes on every changed file, and Pyright reports no errors intabpfn/analytics.E2E smoke tests: fresh processes against a fake gapi, on this version (54 scenarios)
Background thread
Latency
Exceptions (results compared with a run without analytics)
KeyboardInterruptreaches the caller.Concurrency
n_preprocessing_jobs=4, deliver exactly 18 events.multiprocessing.Pooldelivers every worker's events.multiprocessing.Poolfinishes, and its workers log nothing.Unreachable API (events are kept in memory only)
E2E smoke tests on the earlier on-disk version (47 scenarios)
atexiterror when the thread could not start, a race that could re-install the sink after "not enabled", and a macOS crash and pool hang in forked processes.Earlier harness runs, before delivery moved to memory only
Real-estimator break harness (16 scenarios; tests the decorator, unchanged since)
tuning_configgives one fit despite its hidden holdout fits.cross_val_scorelogs 3 fits and 3 predicts, also on 3 threads, and loky workers log their own calls.**kwargsoptions are logged (fixed during testing).KeyboardInterruptpropagate unchanged and are logged as failed.Delivery break harness (33 scenarios, against the earlier on-disk backlog)
collect()(worst 0.3 ms), and events beyond the 10k queue are dropped.stop().stop()beforestart()returns quietly (fixed during testing).stop()twice is harmless.*.jsonamong saved batches no longer blocked the rest (fixed during testing).api_urlwithout a scheme keeps the batch instead of losing it (fixed during testing).exit(),sys.exit(3), an uncaught exception or Ctrl-C delivers every event.os._exit()lose the queued events, asatexitdoes not run (accepted, as in PostHog).multiprocessing.Poolworkers lose their events, as they exit withoutatexit(accepted).event_id.Root cause of the macOS fork crash
🤖 Generated with Claude Code