Skip to content

Latest commit

 

History

History
161 lines (128 loc) · 7.26 KB

File metadata and controls

161 lines (128 loc) · 7.26 KB

Data Schema — Many Phones, One g

The shared format that makes the live session plug-and-play

Schema v1 · draft · 2026-06-17 · CC BY 4.0 · Local Stewardship: U. Warring, Freiburg

Identity card. The one piece of shared infrastructure for the Sensing 2026 consensus-g tutorial (design note: tutorial-consensus-g.md). Two parts: a raw-trace CSV every phone exports, and a result card (YAML) every phone-run, team, and the cohort emits. If every team validates its pipeline against this in the spaced week, live analysis is running a validated pipeline on fresh data — not writing one under time pressure. Declared context: assumes phyphox (RWTH Aachen) or any app that exports time-stamped 3-axis acceleration; SI units throughout.


Why two parts

A raw trace alone lets two teams plot the same swing but not combine their answers: comparability lives in the metadata — which pendulum model, which length convention, whether the large-angle correction was applied and to what. So the schema carries both: the trace and the card that makes the number defensible.

Part 1 — raw-trace CSV (*_trace.csv)

One row per sample, constant rate. Columns are fixed and named; units are in the names. The phase marker is any periodic channel — the swing axis, or |a| — used only to extract the period, never its amplitude (the whole point: timing is the phone's strong axis).

# consensus-g raw trace · schema v1
# t_s    : seconds from recording start, monotonic
# a_x/y/z: acceleration components, m/s^2 (state gravity_included in the result card)
# Constant sample rate; its nominal value is recorded as sample_rate_hz in the result card.
t_s,a_x_ms2,a_y_ms2,a_z_ms2
0.0000,0.0123,-0.0041,9.8071
0.0050,0.0310,-0.0022,9.8112
0.0100,0.0488,0.0007,9.8090
  • Rate: state it; aim ≥ 100 Hz (period timing improves with sample density up to the device floor — see Anchor 2 · noise & Allan).
  • Gravity: either "with g" or "linear/without g" is fine for timing; just declare which in gravity_included.
  • Clock: the phone's own timestamps are the measurement clock — do not resample onto wall-clock.

Part 2 — result card (*_result.yaml)

One card per phone-run; a team emits one team card combining its phones; the cohort emits one cohort card. phase says which scale. Children are listed in combination.inputs, so the whole nested-consensus tree is reconstructable from the cards alone.

schema_version: 1
phase: phone            # phone | team | cohort
team_id: ""
card_id: ""             # unique; referenced by parents' combination.inputs

device:
  make_model: ""        # e.g. "Samsung Galaxy S21"  (teams use >=2 distinct units)
  os: ""
  app: "phyphox"
  app_version: ""
  sensor: "accelerometer"
  sample_rate_hz:        # stated nominal rate
  gravity_included: true # true = "with g"; false = linear acceleration

method:
  type: "pendulum"       # pendulum | freefall  (freefall only if pre-approved wildcard)
  pendulum_model: "simple"   # simple | physical   (required if type=pendulum)
  rig: ""                # how suspended; note shared-rig vs own-rig (device-isolation choice)
  length:
    symbol: "L"          # L (simple) | L_eff (physical, = I_pivot/(m*d) = d + I_cm/(m*d))
    value_m:
    u_m:                 # uncertainty in the length — usually the dominant Type B
    how_measured: ""     # which tape/ruler
    convention: ""       # simple: pivot-to-CoM of bob; physical: I_pivot basis (NOT I_cm)
  amplitude_deg:         # initial angle theta_0
  large_angle_correction:
    applied:             # true | false
    applied_to: ""       # T | T0 | g    (g bias ~ -2x the period correction if ignored)
  damping_treatment: ""  # e.g. "damped-sinusoid fit"
  n_periods:             # how many periods averaged
  fit_method: ""         # zero-crossing | peak-pick | sinusoid-fit

result:
  g_value_ms2:
  u_value_ms2:           # combined standard uncertainty, k=1
  budget:
    type_A_ms2:          # statistical (averaging-limited)
    type_B:              # itemised systematics
      length_ms2:
      amplitude_ms2:
      clock_ms2:
      other_ms2:

combination:             # how children became this card (team/cohort phases)
  rule: ""               # e.g. "inverse-variance weighted mean"; pre-registered before seeing others
  inputs: []             # card_ids combined

environment:
  room: ""
  floor:
  coords_etrs89: ""      # for the BKG lookup
  height_dhhn2016_m:
  temperature_c:
  reference_clock: ""

provenance:
  recorded_by: ""
  date: ""               # ISO 8601
  raw_trace_file: ""     # link to the matching *_trace.csv

Minimum comparability set

A card may be combined into a parent only if these are present and non-empty: result.g_value_ms2, result.u_value_ms2, the type_A / type_B split, method.type, method.pendulum_model, method.length (symbol + value + u + convention), large_angle_correction (applied + applied_to), n_periods, and environment.room/floor. A card missing any of these is plottable but not combinable — surface it, don't silently drop it.

The one hard rule: the external value is a comparator, not an input

The BKG geodetic value is ~10²× more precise than the phone consensus, so inverse-variance weighting would let it silently dominate and the exercise would collapse to "trust BKG." Therefore:

  • It is recorded in its own object (below), never in any combination.inputs.
  • The cohort reports the phone-only consensus first, then the comparison.
external_reference:
  source: "BKG online gravity calculator"   # gibs.bkg.bund.de/geoid/gscomp.php?p=s
  g_value_ms2:                               # run for the test room (ETRS89 + DHHN2016 height)
  stated_accuracy_mGal: 2                    # <2 mGal in Germany (<7 in rugged topography)
  retrieved: ""                              # ISO date
comparison:
  phone_consensus_g_ms2:
  phone_consensus_u_ms2:
  delta_ms2:                                 # phone_consensus - external
  covers_external:                           # does the phone interval contain the BKG value?

How it threads the run sheet

  • Prep: validate the pipeline so a *_trace.csv → *_result.yaml round-trip is automatic.
  • 0–10 min: post the team card with the budget and combination.rule filled — pre-registration (revisable once, with spoken justification).
  • 30–45 min: fill result from live data; the cohort builds its card from the team cards.
  • 55–60 min: fill external_reference + comparison; this is the debrief reveal.

The cohort comparison.covers_external is the headline of the logbook entry — not "how close," but "did our honest interval cover the independent value?"

Provenance and changelog

Drafted 2026-06-17 alongside tutorial-consensus-g.md v0.5. External-anchor URL and accuracy verified on the web 2026-06-17 (BKG / PTB).

  • v1 (2026-06-17) — first draft: raw-trace CSV, result-card YAML, minimum comparability set, external-value-as-comparator rule, run-sheet threading.

Sensing 2026 · data schema · University of Freiburg (group of T. Schaetz). Text © 2026 U. Warring · CC BY 4.0 · schema v1.