You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Latency is not comparable across the sections above. Attention cost grows faster than pixel count -- measured on depth_anything_ac, 1.35x the pixels cost 1.70x the time -- so dividing by resolution does not rescue the comparison.
Profile native
Input 1536x1536 (single model)
model
variant
encoder
precision
mean ms
p50
p90
fps
fp32
moved
output
depth_pro
single
fixed
fp16
242.12
242.17
243.15
4.13
0.6%
0.4%
metric
depth_pro -- 1536x1536 is upstream fixed; no 518 bench profile exists
Input 616x1064 (single model)
model
variant
encoder
precision
mean ms
p50
p90
fps
fp32
moved
output
metric3d_v2
single
vits
fp32
62.39
62.39
62.65
16.03
100.0%
4.6%
canonical (not metres; needs a real focal length)
Latency is not comparable across the sections above. Attention cost grows faster than pixel count -- measured on depth_anything_ac, 1.35x the pixels cost 1.70x the time -- so dividing by resolution does not rescue the comparison.
Caveats
These numbers are valid as latency. The output carries a condition recorded in docs/model_contracts.md.
depth_anything_ac -- output_form is inverse_depth: larger means nearer. Measured, not assumed -- the prediction correlates -0.927 with true depth over five DIODE images. It decides whether the scale+shift fit is affine in the space the model was trained in
depth_anything_v3 -- the sky refinement pass is not in the graph: it samples randomly and cannot be traced (core/export_compat.no_mono_sky_postprocess)
depth_pro -- 1536x1536 is upstream's fixed size; there is no 518 bench profile
distill_any_depth -- output_form is inverse_depth: larger means nearer. Measured, not assumed -- the prediction correlates -0.938 with true depth over five DIODE images. It decides whether the scale+shift fit is affine in the space the model was trained in
metric3d_v2 -- output is canonical depth; metres need real_focal * scale / 1000, and that focal length has to come from the dataset -- an arbitrary photograph does not carry one; built at fp32 while every other model is fp16
metric_anything -- 388x518 assumes a 4:3 source, matching moge_2 (D13); post-process calls MoGe's recover_focal_shift, so trte needs the pinned utils3d commit (D15)
moge_2 -- 388x518 assumes a 4:3 source; post-process calls MoGe's recover_focal_shift, so trte needs the pinned utils3d commit (D15)
streamvggt -- cartesian_prod is replaced at export time; TorchScript has no symbolic for it (core/export_compat.no_cartesian_prod)
tr2m -- the only two-input model here: image plus a [1,1,768] CLIP text embedding. CLIP is not in the engine, so the prompt is fixed at export time and its embedding is committed as text_features.npy; 434x560, not the 518x518 the rest use, so a frames-per-second comparison is between different amounts of work; upstream has NO LICENSE file at all, and its pos_embed.py carries 'Copyright (C) 2022-present Naver Corporation. All rights reserved. Licensed under CC BY-NC-SA 4.0 (non-commercial use only)'. Read from the repository 2026-08-14. The clone stays out of this repository and nothing is vendored, which is what keeps that away from this repository's MIT; metric depth is produced from a relative model by predicting a per-pixel scale and shift, so its errors are those of Depth Anything plus those of the rescaling
unidepth_v2 -- 672x896 because that is what the model's own resize rule returns for a 4:3 source -- aspect preserved, pixel count into [200000, 600000], multiple of 14. Running it at 518x518 told it that was the native resolution, it inferred the focal length from that, and the metric depth came out 3.1x upstream. That was the cost of the override, not a conversion defect. One engine covers a fixed-resolution dataset; a different source aspect ratio needs a different build; float64 in the exported graph is rewritten to float32 after export (core/onnx_tools.demote_float64)
unik3d -- 672x896 because that is what the model's own resize rule returns for a 4:3 source -- aspect preserved, pixel count into [200000, 600000], multiple of 14. Running it at 518x518 told it that was the native resolution, it inferred the focal length from that, and the metric depth came out 3.1x upstream. That was the cost of the override, not a conversion defect. One engine covers a fixed-resolution dataset; a different source aspect ratio needs a different build; exported with antialias disabled; TensorRT cannot parse antialias=1
vggt -- cartesian_prod is replaced at export time; TorchScript has no symbolic for it (core/export_compat.no_cartesian_prod)
zipdepth -- the only convolutional model here at 6.1M parameters; every other row is a ViT or a transformer geometry network, so it exercises TensorRT's convolution path rather than its attention path; 384x512 fixes upstream's rule (short side 384, both sides a multiple of 32) at the 4:3 of data/example.jpg. A different aspect ratio needs a different engine; preprocessing is /255 only. There is no ImageNet mean/std here, unlike every other model in this repository; output is inverse depth, so larger means nearer. Scored with a scale-and-shift fit on disparity, which is the form it is already in; trained on pseudo-labels from Depth Anything V2 Large, whose weights are CC BY-NC 4.0. Whether that reaches a distilled student is unsettled; the code and checkpoints here are MIT; exported in the shared trte environment rather than one of its own: onnx_export.py puts the clone on sys.path instead of pip-installing it, and torch 2.11 there already satisfies upstream's torch>=2.4. Verified 2026-08-14 -- the export ran clean with no missing state_dict keys; output_form is inverse_depth: larger means nearer. Measured, not assumed -- the prediction correlates -0.892 with true depth over five DIODE images. It decides whether the scale+shift fit is affine in the space the model was trained in