Skip to content

Latest commit

 

History

History
90 lines (63 loc) · 9.21 KB

File metadata and controls

90 lines (63 loc) · 9.21 KB

Model comparison

Generated by compare.py from reports/bench/*.json. Do not edit by hand -- rerun the script.

  • GPU: NVIDIA GeForce RTX 3080
  • TensorRT: 10.16.1.11
  • Measured: 2026-08-13 to 2026-08-14
  • Runs: 14

Mixed precision. Most rows are fp16, but fp32: metric3d_v2. Compare within a precision, not across.

Profile bench

Input 518x518

model variant encoder precision mean ms p50 p90 fps fp32 moved output
depth_anything_ac single vits fp16 4.38 4.37 4.41 228.47 1.5% 1.9% relative
depth_anything_v2 single - fp16 4.31 4.31 4.35 232.11 1.8% 1.9% metric (hypersim) by default
depth_anything_v3 single DA3Metric-Large fp16 19.90 19.89 19.95 50.26 0.6% 0.5% metric + sky mask
distill_any_depth single small fp16 4.32 4.32 4.35 231.47 1.7% 1.9% relative
streamvggt single fixed fp16 52.83 52.82 52.90 18.93 4.1% 0.1% geometry (scale unknown)
vggt single fixed fp16 52.58 52.58 52.64 19.02 4.1% 0.1% geometry (scale unknown)
  • depth_anything_v2 -- encoder=vits, weights=metric_hypersim

Input 388x518

model variant encoder precision mean ms p50 p90 fps fp32 moved output
metric_anything single fixed fp16 67.00 66.99 67.23 14.93 1.3% 14.7% point map + metric_scale
moge_2 single vits fp16 19.73 19.72 19.77 50.68 1.8% 32.2% point map + metric_scale

Input 672x896

model variant encoder precision mean ms p50 p90 fps fp32 moved output
unidepth_v2 single vits fp16 17.82 17.82 17.88 56.12 0.5% 9.4% point map + intrinsics
unik3d single vits fp16 18.35 18.34 18.42 54.51 1.1% 9.0% point map + intrinsics

Input 434x560 (single model)

model variant encoder precision mean ms p50 p90 fps fp32 moved output
tr2m single vits fp16 20.08 20.08 20.16 49.80 0.5% 0.7% metric (text-conditioned)

Input 384x512 (single model)

model variant encoder precision mean ms p50 p90 fps fp32 moved output
zipdepth single base fp16 2.77 2.68 3.11 361.40 11.6% 8.6% relative (inverse depth)

Latency is not comparable across the sections above. Attention cost grows faster than pixel count -- measured on depth_anything_ac, 1.35x the pixels cost 1.70x the time -- so dividing by resolution does not rescue the comparison.

Profile native

Input 1536x1536 (single model)

model variant encoder precision mean ms p50 p90 fps fp32 moved output
depth_pro single fixed fp16 242.12 242.17 243.15 4.13 0.6% 0.4% metric
  • depth_pro -- 1536x1536 is upstream fixed; no 518 bench profile exists

Input 616x1064 (single model)

model variant encoder precision mean ms p50 p90 fps fp32 moved output
metric3d_v2 single vits fp32 62.39 62.39 62.65 16.03 100.0% 4.6% canonical (not metres; needs a real focal length)

Latency is not comparable across the sections above. Attention cost grows faster than pixel count -- measured on depth_anything_ac, 1.35x the pixels cost 1.70x the time -- so dividing by resolution does not rescue the comparison.

Caveats

These numbers are valid as latency. The output carries a condition recorded in docs/model_contracts.md.

  • depth_anything_ac -- output_form is inverse_depth: larger means nearer. Measured, not assumed -- the prediction correlates -0.927 with true depth over five DIODE images. It decides whether the scale+shift fit is affine in the space the model was trained in
  • depth_anything_v3 -- the sky refinement pass is not in the graph: it samples randomly and cannot be traced (core/export_compat.no_mono_sky_postprocess)
  • depth_pro -- 1536x1536 is upstream's fixed size; there is no 518 bench profile
  • distill_any_depth -- output_form is inverse_depth: larger means nearer. Measured, not assumed -- the prediction correlates -0.938 with true depth over five DIODE images. It decides whether the scale+shift fit is affine in the space the model was trained in
  • metric3d_v2 -- output is canonical depth; metres need real_focal * scale / 1000, and that focal length has to come from the dataset -- an arbitrary photograph does not carry one; built at fp32 while every other model is fp16
  • metric_anything -- 388x518 assumes a 4:3 source, matching moge_2 (D13); post-process calls MoGe's recover_focal_shift, so trte needs the pinned utils3d commit (D15)
  • moge_2 -- 388x518 assumes a 4:3 source; post-process calls MoGe's recover_focal_shift, so trte needs the pinned utils3d commit (D15)
  • streamvggt -- cartesian_prod is replaced at export time; TorchScript has no symbolic for it (core/export_compat.no_cartesian_prod)
  • tr2m -- the only two-input model here: image plus a [1,1,768] CLIP text embedding. CLIP is not in the engine, so the prompt is fixed at export time and its embedding is committed as text_features.npy; 434x560, not the 518x518 the rest use, so a frames-per-second comparison is between different amounts of work; upstream has NO LICENSE file at all, and its pos_embed.py carries 'Copyright (C) 2022-present Naver Corporation. All rights reserved. Licensed under CC BY-NC-SA 4.0 (non-commercial use only)'. Read from the repository 2026-08-14. The clone stays out of this repository and nothing is vendored, which is what keeps that away from this repository's MIT; metric depth is produced from a relative model by predicting a per-pixel scale and shift, so its errors are those of Depth Anything plus those of the rescaling
  • unidepth_v2 -- 672x896 because that is what the model's own resize rule returns for a 4:3 source -- aspect preserved, pixel count into [200000, 600000], multiple of 14. Running it at 518x518 told it that was the native resolution, it inferred the focal length from that, and the metric depth came out 3.1x upstream. That was the cost of the override, not a conversion defect. One engine covers a fixed-resolution dataset; a different source aspect ratio needs a different build; float64 in the exported graph is rewritten to float32 after export (core/onnx_tools.demote_float64)
  • unik3d -- 672x896 because that is what the model's own resize rule returns for a 4:3 source -- aspect preserved, pixel count into [200000, 600000], multiple of 14. Running it at 518x518 told it that was the native resolution, it inferred the focal length from that, and the metric depth came out 3.1x upstream. That was the cost of the override, not a conversion defect. One engine covers a fixed-resolution dataset; a different source aspect ratio needs a different build; exported with antialias disabled; TensorRT cannot parse antialias=1
  • vggt -- cartesian_prod is replaced at export time; TorchScript has no symbolic for it (core/export_compat.no_cartesian_prod)
  • zipdepth -- the only convolutional model here at 6.1M parameters; every other row is a ViT or a transformer geometry network, so it exercises TensorRT's convolution path rather than its attention path; 384x512 fixes upstream's rule (short side 384, both sides a multiple of 32) at the 4:3 of data/example.jpg. A different aspect ratio needs a different engine; preprocessing is /255 only. There is no ImageNet mean/std here, unlike every other model in this repository; output is inverse depth, so larger means nearer. Scored with a scale-and-shift fit on disparity, which is the form it is already in; trained on pseudo-labels from Depth Anything V2 Large, whose weights are CC BY-NC 4.0. Whether that reaches a distilled student is unsettled; the code and checkpoints here are MIT; exported in the shared trte environment rather than one of its own: onnx_export.py puts the clone on sys.path instead of pip-installing it, and torch 2.11 there already satisfies upstream's torch>=2.4. Verified 2026-08-14 -- the export ran clean with no missing state_dict keys; output_form is inverse_depth: larger means nearer. Measured, not assumed -- the prediction correlates -0.892 with true depth over five DIODE images. It decides whether the scale+shift fit is affine in the space the model was trained in