This document describes all manual changes applied on top of the MCC-generated Harmony 3 project for the ATSAME54P20A + LAN865x 10BASE-T1S demo. The goal is to enable PTP (IEEE 1588) hardware timestamping with sub-microsecond synchronisation accuracy over 10BASE-T1S.
Sync quality (current HEAD — freeze candidate, measured 2026-04-24): Cross-board PD10 drift MAD 13.6 µs, slope −0.07 ppm over 10 s (canonical test tools/ptp-analysis/sync-tests/pd10_sync_before_after_test.py, run
20260424_084628). All PASS gates met with wide margin (|slope| < 5 ppm,MAD < 50 µs). See §5.12 for details.
Navigation
- Full documentation index: documentation/README.md
- Developer tools (flash / tests / analysis): tools/README.md
- Risks & open questions: RISKS.md
- Cleanup history: CLEANUP_PLAN.md
python setup_flasher.py :: once — detect the two debugger COM ports
python flash.py :: program both boards with the checked-in default.hex
cd tools\test-harness && python smoke_test.py :: optional: 58-check regression
:: To rebuild the firmware from source:
python setup_compiler.py :: once — pick an installed XC32 version
build.bat :: incremental build (build.bat rebuild for clean)
python flash.py :: re-flash (default.hex was overwritten by the build)flash.py with no argument always programs
apps/tcpip_iperf_lan865x/firmware/tcpip_iperf_lan865x.X/out/tcpip_iperf_lan865x/default.hex
— that file is checked in, so a fresh clone is immediately flash-ready, and
a local build.bat run overwrites it in place, so the next flash.py
picks up your new build automatically. Full walkthrough:
§How To Reproduce.
Top-level scripts (setup_*.py, build.bat, build_summary.py,
flash.py, mdb_flash.py, check_serial_tk.pyw) live at the repo root;
test / analysis / Saleae scripts live under tools/. The MPLAB X
project itself is untouched at
apps/tcpip_iperf_lan865x/firmware/tcpip_iperf_lan865x.X/.
For the full PTP implementation reference — state machine pseudocode, IEEE 1588-2008 compliance verification, Mermaid flow diagrams, servo design, register reference, and the automated regression test — see:
documentation/ptp/implementation.md
For the companion Software-NTP module — application-layer UDP time-sync using
SW timestamps from PTP_CLOCK_GetTime_ns(), a two-phase test that compares the
SW-NTP jitter floor with and without HW-PTP disciplining, and a discussion of
how precisely the software clock can be used from application code — see:
documentation/ptp/ntp_reference.md
For the tfuture module — coordinated single-shot firing at an absolute PTP_CLOCK time across two HW-PTP-synchronised boards (self_jitter < 50 µs median, inter-board coincidence within ±30 µs), bias-investigation history that uncovered and fixed two latent bugs in the PTP_CLOCK drift filter, and a by-product that reports the ppm deviations of all four crystals in the system (2× SAME54 + 2× LAN8651) — see:
documentation/features/tfuture.md
For the standalone demo — two-board self-contained PTP sync demo
driven entirely by the on-board SW1/SW2 buttons and LED1/LED2 indicators
(no PC required during the demo). Branches ptp-standlone-demo and
iperf-payload-test add a decimated 1 Hz visible blink that visually
drifts apart before PTP is activated and snaps back into lock-step once
the SW1/SW2 role-selection is made, plus an iperf TCP server/client
toggle on the opposite button of each role that exchanges payload over
the synchronised link:
documentation/features/standalone_demo.md
For the PTP_CLOCK drift filter — adaptive single-pole IIR design that breaks the usual settle-vs-jitter trade-off (warm-up α=1 → steady-state α=1/N_max), runtime CLI tuning, fixed-vs-adaptive measurement comparison on the cyclic_fire cross-board edges (~7 µs MAD vs ~25 µs MAD), and guidance on choosing N_max for different operating conditions — see:
documentation/ptp/drift_filter.md
For the exception dump + watchdog + find_exception.py crash-diagnostics
subsystem — Cortex-M4 fault handler trampolines + register dump over
SERCOM1, SAM E54 WDT with Early-Warning catching silent hangs,
test_exception CLI for verifying the dump path, and the Python decoder
that maps the dumped PC back to the C source line — see:
documentation/hardware/exception_dump.md
This repository is a fork of the Microchip Harmony 3 net_10base_t1s package:
https://github.com/Microchip-MPLAB-Harmony/net_10base_t1s.git Base commit:
586ffc15708fcc2c02b182967d872837a15f69f7(tag: v1.4.3)
The PTP implementation added on top is derived from Microchip Application Note AN1847 (Precision Time Protocol over 10BASE-T1S) and its accompanying open-source reference project, which is itself a fork of:
The AN1847 reference project targets a different Harmony 3 demo application
(t1s_100baset_bridge). The work described here ports and adapts that
implementation to the tcpip_iperf_lan865x project, resolves two bugs found
during testing (see §2), and adds the ptp_clock.c software wallclock which
is not present in the original.
- How To Reproduce
- Overview
- 1. PTP Implementation
- 3. PTP Software Clock
- 4. Further Firmware Changes
- 5. Test Scripts & Validation
- How the Tests Control the Firmware
- 5.1 Baseline: TC0 Timer Consistency
- 5.2 PTP On/Off Resilience
- 5.3 Role-Swap Validation
- 5.4 PTP Sync Before/After
- 5.5 PTP Reproducibility
- 5.6 IEEE 1588 Compliance + Convergence Regression
- 5.7 PTP Drift Compensation + Loop-Stats Instrumentation
- 5.8 Before/After (mux-based)
- 5.9 Hardware-Timestamp Offset Capture
- 5.10 Software NTP vs HW-PTP Comparison
- 5.11 tfuture Coordinated Firing + Crystal Deviation Analysis
- 5.12 Cyclic GPIO Toggle — cyclic_fire
- 5.13 Standalone PD10 rectangle generator — blink
- 6. Hardware & Build Setup
- 7. Reinforcement Learning — Coding with AI
- 8. Python Dependency Management
- 9. PTP Implementation — In-Depth Analysis
| Area | Changed / New File(s) | Purpose |
|---|---|---|
| PTP hardware timestamping | drv_lan865x_api.c, drv_lan865x.h, tc6.c |
TX/RX timestamps for PTP Sync/FollowUp |
| PTP state machines | ptp_gm_task.c/.h, ptp_fol_task.c/.h, filters.c/.h, ptp_ts_ipc.h |
Grandmaster + Follower roles |
| PTP software clock | ptp_clock.c/.h |
ns-resolution wallclock via TC0 |
| Application integration | app.c, app.h |
PTP services, packet handler, CLI |
| Bug fix — TX timestamp | drv_lan865x_api.c |
DELAY_UNLOCK_EXT 100 ms → 5 ms; fixes missed timestamps |
| Bug fix — role-swap | ptp_gm_task.c, ptp_fol_task.c/.h |
−3.13 ms stuck offset after GM↔FOL swap |
| IEEE 1588 compliance fixes | ptp_gm_task.c, ptp_fol_task.c |
5 fixes: twoStepFlag, tsmt bytes, sequence-ID verification, TXMPATL pattern |
| MAC randomisation | initialization.c |
Unique MAC addresses via hardware TRNG |
| LAN865x register CLI | app.c |
lan_read / lan_write without a debugger |
| Build tooling | build.bat, setup_compiler.py, setup_flasher.py, setup_debug.py, flash.py, build_summary.py, user.cmake |
Reproducible one-command builds; setup_debug.py fixes a DFP tool-pack bug that prevents VS Code debugging |
LAN865x SPI Hardware
|
| TX Sync frame → TTSCAA bit set in STATUS0
↓
_OnStatus0() [drv_lan865x_api.c]
→ saves STATUS0 bits 8-10 into drvTsCaptureStatus0[]
→ W1C-clears STATUS0
|
↓
DRV_LAN865X_GetAndClearTsCapture(0) [called by PTP_GM_Service]
→ GM reads TX timestamp from TTSCAH/AL registers
→ sends FollowUp with corrected timestamp
-- RX path (primary) --
TC6_CB_OnRxEthernetPacket() [drv_lan865x_api.c]
→ copies frame + rxTimestamp into g_ptp_raw_rx (global, ptp_ts_ipc.h)
→ sets g_ptp_raw_rx.pending = true
|
↓
APP_STATE_IDLE [app.c]
→ checks g_ptp_raw_rx.pending first
→ PTP_FOL_OnFrame(data, length, rxTimestamp)
→ Follower servo computes offset and adjusts local clock
-- TCP/IP stack path (frame suppression) --
TC6_CB_OnRxEthernetPacket() → TCP/IP stack → pktEth0Handler() [app.c]
→ EtherType 0x88F7: TCPIP_PKT_PacketAcknowledge(TCPIP_MAC_PKT_ACK_RX_OK)
→ returns true (consumed) — IP stack does not process PTP frames
→ no frame copy; driver path (g_ptp_raw_rx) is the single source of truth
The following steps take you from a fresh git clone to live PTP output on two
boards in about 10 minutes. All commands run from the repo root
(net_10base_t1s/) — the top-level driver scripts (setup_*.py, build.bat,
flash.py, mdb_flash.py) were moved to the root in 2026-04 so the workflow
no longer requires cd-ing into apps/.../tcpip_iperf_lan865x.X/.
| Requirement | Notes |
|---|---|
| Hardware | 2× ATSAME54-Curiosity-Ultra + LAN865x click board, connected via T1S bus |
| MPLAB XC32 | v4.60 or v5.x, installed under C:\Program Files\Microchip\xc32\ (only needed for the build path) |
| CMake ≥ 4.1 + Ninja | On PATH (only needed for the build path) |
| MPLAB X IDE / MDB | Required by flash.py for programming — the mdb CLI is bundled with MPLAB X |
| Python 3.9+ | pip install pyserial for the serial test scripts |
| Terminal emulator | 115200 8N1 (PuTTY, Tera Term, etc.) — only for the live CLI console |
git clone https://github.com/zabooh/net_10base_t1s.git
cd net_10base_t1sAll further commands assume you are at this repo root (not inside apps/).
Three interactive scripts, run once per machine. Each saves its result
to a git-ignored .config file, so you never have to repeat them on later
sessions.
python setup_flasher.py :: detect + assign Board 1 (GM) / Board 2 (FOL) debuggers
python setup_compiler.py :: pick an installed XC32 version (ONLY if you will build)
python setup_debug.py :: patch SAME54_DFP tool-pack (ONLY if you will debug in VS Code)setup_flasher.py— both boards must be connected via USB (EDBG port) when you run this. The script enumerates all plugged-in EDBG debuggers and lets you assign which serial is Board 1 vs Board 2. Result written tosetup_flasher.config.setup_compiler.py— skip this if you are only flashing the pre-built HEX. Needed before the firstbuild.batcall.setup_debug.py— skip this unless you plan to step through code in VS Code withcortex-debug. Does not affect flashing or the command-line build.
After cloning, the latest known-good firmware HEX is already checked in at
apps/tcpip_iperf_lan865x/firmware/tcpip_iperf_lan865x.X/out/tcpip_iperf_lan865x/default.hex.
flash.py programs exactly that file by default — no compiler or CMake
needed on your machine:
python flash.py :: programs both boards with default.hex
python flash.py --board1-only :: only Board 1 (GM)
python flash.py --board2-only :: only Board 2 (FOL)
python flash.py --hex <path> :: alternative HEX (absolute or relative to repo root)Both boards run identical firmware. The role (Grandmaster / Follower) is assigned at runtime via the
ptp_mode master/ptp_mode followerCLI — no separate images.
build.bat :: incremental build — only changed files
build.bat rebuild :: clean + full rebuild
build.bat clean :: remove build artifactsA successful build overwrites out/.../default.hex in place with a fresh
image (and also produces a timestamped copy under out/.../image/). The
next python flash.py therefore automatically picks up your new build — no
path argument needed:
build.bat && python flash.py :: build then flash, one lineA build summary (flash / RAM usage, heap, active interrupt handlers, HEX path) is printed automatically after every successful build — see §6.2 for the full description.
Once both boards are flashed, run the broad functional regression guard to confirm the build is healthy. Takes ~60 s and exercises every CLI command, PTP FINE lock, the tfuture single-shot firing, the cyclic_fire ISR path, SW-NTP end-to-end exchange, and a cross-board FOL/GM offset bracketing:
cd tools\test-harness
python smoke_test.pyUseful flags:
| Flag | Default | Effect |
|---|---|---|
--gm-port, --fol-port |
COM8, COM10 |
serial ports of GM / FOL |
--no-reset |
off | skip Phase 1 boot + PTP FINE (assume boards already running) |
--abort-on-fail |
off | stop on the first FAIL for fast-fail CI use |
--verbose |
off | log every serial round-trip |
--settle-s N |
5.0 |
extra seconds after PTP FINE before Phase 2 starts |
Expected happy-path ending:
======================================================================
Summary: 58 PASS / 0 FAIL (total 58)
======================================================================
======================================================================
Post-test: reset + PTP sync (leave boards ready for next use)
======================================================================
...
FOL reached PTP FINE in 2.1 s — GM (master) + FOL (follower) in sync,
ready for next use.
The post-test teardown is automatic: after the summary prints, both boards are reset, IPs are reconfigured, PTP is restarted (GM = master, FOL = follower), and the follower is waited for FINE — so the boards are left in a clean, ready-to-use synced state that you can immediately attach a Saleae to, or drive via the CLI, without another reset.
Exit code: 0 on all-PASS, 1 on any FAIL. A run log is written to
smoke_test_<YYYYmmdd_HHMMSS>.log in the tools/test-harness/ directory.
Open two serial terminal windows (115200 8N1, no flow control):
- Board 1 — Grandmaster (COM8 by default)
- Board 2 — Follower (COM10 by default)
After reset both boards print the Harmony boot banner and then stay idle. Activate PTP from each terminal:
Board 1 — Grandmaster (verbose mode):
> ptp_mode master v
Board 2 — Follower (verbose mode):
> ptp_mode follower v
Expected Board 2 output — the servo steps through these states:
PTP MATCHFREQ offset=...
PTP HARDSYNC offset=...
PTP COARSE offset=...
PTP FINE offset=...
[FOL] FINE t1=00:00:05.123456789 t2=00:00:05.123514600 off= +57 ns
[FOL] FINE t1=00:00:05.248334810 t2=00:00:05.248392100 off= +55 ns
Once the follower prints FINE continuously, synchronisation is
established. Offsets below ±200 ns are typical for the initial lock;
steady-state offsets are usually below ±1 µs. Check the servo state at
any time:
> ptp_status
PTP support was ported from the reference project
C:\work\ptp\AN1847\t1s_100baset_bridge\ to this project
(see also Origin above).
The implementation supports both Grandmaster (GM) and Follower (FOL)
roles, switchable at runtime via CLI.
| File | Description |
|---|---|
src/ptp_gm_task.c/.h |
PTP Grandmaster state machine. Sends Sync + FollowUp frames at a configurable interval, arms the LAN865x TX-Match hardware for TX timestamp capture. |
src/ptp_fol_task.c/.h |
PTP Follower state machine. Receives Sync/FollowUp, computes clock offset with FIR low-pass filter, and slaves the local time. |
src/ptp_clock.c/.h |
TC0-based nanosecond software wallclock. See §3. |
src/ptp_ts_ipc.h |
Shared IPC header: PTP_RxTimestampEntry_t struct + g_ptp_rx_ts extern declaration. |
src/filters.c/.h |
FIR low-pass filter and exponential low-pass filter used by the Follower servo. |
src/ptp_offset_trace.c/.h |
1024-entry ring buffer for hardware-timestamp PTP offsets. Used by ptp_offset_capture.py (§5.9). |
src/sw_ntp.c/.h |
Minimal software-NTP (UDP) master + follower using SW timestamps from PTP_CLOCK_GetTime_ns(). Measurement-only; does not discipline the clock. See documentation/ptp/ntp_reference.md. |
src/sw_ntp_offset_trace.c/.h |
1024-entry int64 ring buffer for SW-NTP offsets. Used by sw_ntp_vs_ptp_test.py (§5.10). |
src/tfuture.c/.h |
Coordinated single-shot firing at an absolute PTP_CLOCK time. Hybrid precision (main-loop poll + runtime-configurable tight-spin). 256-entry ring buffer. Post-fire callback hook for periodic re-arming. See documentation/features/tfuture.md and §5.11. |
src/cyclic_fire.c/.h |
PTP-synchronous periodic GPIO toggle on PD10 at a configurable period_us (default 1000 µs → 1 kHz rectangle). Two output patterns: SQUARE (50/50 toggle, for rate/phase measurement) and MARKER (1-high + 4-low pulse, for visual "who fires first?"). Uses the tfuture callback hook. See §5.12. |
src/cyclic_fire_cli.c/.h |
CLI wrapper for cyclic_fire — registers cyclic_start / cyclic_start_marker / cyclic_start_free / cyclic_stop / cyclic_status. |
src/pd10_blink.c/.h |
Standalone main-loop rectangle generator on PD10 at a configurable frequency. Independent of PTP and tfuture; pure SYS_TIME_Counter64Get service. See §5.13. |
src/pd10_blink_cli.c/.h |
CLI wrapper for pd10_blink — registers the blink command. |
src/lan_regs_cli.c/.h |
CLI adapter + async state machine for lan_read / lan_write LAN865x register access (extracted from app.c). |
src/ptp_cli.c/.h |
CLI adapter for all 14 PTP / clock / offset-trace commands (extracted from app.c). |
src/sw_ntp_cli.c/.h |
CLI adapter for all 6 SW-NTP commands + IP parser (extracted from app.c). |
src/tfuture_cli.c/.h |
CLI adapter for all 7 tfuture commands (extracted from app.c). |
src/loop_stats_cli.c/.h |
CLI adapter for the loop_stats command (extracted from app.c). |
src/ptp_rx.c/.h |
PTP 0x88F7 frame filter + dispatcher between Harmony TCP/IP stack and the PTP tasks (extracted from app.c). |
Three source files added to the CMake build in cmake/.../CMakeLists.txt:
"${CMAKE_CURRENT_SOURCE_DIR}/../../../../src/filters.c"
"${CMAKE_CURRENT_SOURCE_DIR}/../../../../src/ptp_fol_task.c"
"${CMAKE_CURRENT_SOURCE_DIR}/../../../../src/ptp_gm_task.c"static volatile uint32_t drvTsCaptureStatus0[DRV_LAN865X_INSTANCES_NUMBER];Shadow register that saves STATUS0 bits 8–10 (TTSCAA/B/C) in _OnStatus0()
before the W1C-clear. Read and atomically cleared by
DRV_LAN865X_GetAndClearTsCapture(). This avoids the race condition where the
driver clears TTSCAA before the GM state machine can read it.
typedef struct { uint64_t rxTimestamp; bool valid; } PTP_RxTimestampEntry_t;
volatile PTP_RxTimestampEntry_t g_ptp_rx_ts = {0u, false};Defined at the top of TC6_CB_OnRxEthernetPacket(). The callback now saves the
hardware RX timestamp into this struct when rxTimestamp != NULL. The
application reads it in pktEth0Handler() when a PTP frame (EtherType 0x88F7)
arrives.
The memory map was updated to enable PTP timestamp hardware:
| Register | Address | Old Value | New Value | Comment |
|---|---|---|---|---|
| TXMPATH | 0x00040041 | (not present) | 0x0088 | EtherType high byte 0x88 |
| TXMPATL | 0x00040042 | (not present) | 0xF700 | EtherType low 0xF7 + PTP messageType 0x00 (Sync, transportSpecific=0) |
| TXMMSKH | 0x00040043 | 0x00FF | 0x0000 | No masking — exact match |
| TXMMSKL | 0x00040044 | 0xFFFF | 0x0000 | No masking |
| TXMLOC | 0x00040045 | 0x0000 | 0x001E | Byte offset 30 (from Microchip PTP demo) |
| TXMCTL | 0x00040040 | 0x0002 | 0x0000 | Disabled at startup; armed per-Sync |
| IMASK0 | 0x0000000C | 0x0100 | 0x0000 | All interrupts unmasked (incl. TTSCAA bit 8) |
| DEEP_SLEEP_CTRL_1 | 0x00040081 | 0x0080 | 0x00E0 | Updated per reference |
| (removed) | 0x000400E0 | 0xC000 | — | Moved to _InitConfig case 46 as PADCTRL RMW |
// Case 46: PADCTRL RMW — enables TX timestamp pad output
TC6_ReadModifyWriteRegister(tc, 0x000A0088u, 0x00000100u, 0x00000300u, ...);
// Case 47: PPSCTL — enables PPS clock for TSU counter
TC6_WriteRegister(tc, 0x000A0239u, 0x0000007Du, ...);Previously case 46 wrote 0xC000 to 0x000400E0. Replaced with the PADCTRL
RMW required for TX hardware timestamping.
regVal = 0x9026u;
regVal |= 0x80u; // FTSE: Frame Timestamp Enable
regVal |= 0x40u; // FTSS: 64-bit timestampsEnables frame-level timestamping in CONFIG0. Required for TTSCAA TX capture and the TC6 driver's RTSA 8-byte timestamp stripping on RX.
if (0u != (value & 0x0F00u)) {
SYS_CONSOLE_PRINT("[DBG] _OnStatus0: 0x%08lX\r\n", (unsigned long)value);
}
if (0u != (value & 0x0700u)) {
for (i = 0u; i < DRV_LAN865X_INSTANCES_NUMBER; i++) {
if (pDrvInst == &drvLAN865XDrvInst[i]) {
drvTsCaptureStatus0[i] |= (value & 0x0700u);
break;
}
}
}
// ... then W1C WriteRegister ...STATUS0 bits 8–10 are saved into drvTsCaptureStatus0[] before the
Write-1-to-Clear operation.
| Function | Description |
|---|---|
DRV_LAN865X_SendRawEthFrame(idx, pBuf, len, tsc, cb, pTag) |
Sends a raw Ethernet frame via TC6. TSC flag 0x01 = Capture A for Sync, 0x00 = no capture. |
DRV_LAN865X_IsReady(idx) |
Returns true when the driver instance is fully initialised. Used to detect Loss-of-Framing recovery. |
DRV_LAN865X_GetAndClearTsCapture(idx) |
Atomically reads and clears drvTsCaptureStatus0[idx]. Called by the GM state machine to retrieve TTSCAA/B/C bits. |
bool DRV_LAN865X_SendRawEthFrame(uint8_t idx, const uint8_t *pBuf,
uint16_t len, uint8_t tsc,
DRV_LAN865X_RawTxCallback_t cb, void *pTag);
bool DRV_LAN865X_IsReady(uint8_t idx);
uint32_t DRV_LAN865X_GetAndClearTsCapture(uint8_t idx);Added after the existing DRV_LAN865X_ReadModifyWriteRegister declaration:
DRV_LAN865X_RawTxCallback_ttypedef- Declarations for
DRV_LAN865X_SendRawEthFrame(),DRV_LAN865X_IsReady(),DRV_LAN865X_GetAndClearTsCapture()
Two changes vs. the reference project (t1s_100baset_bridge) to support correct
PTP TX timestamp capture indexing:
1. Added include:
#include "driver/lan865x/src/dynamic/drv_lan865x_local.h"2. TC6_Init — instance slot selection rewritten:
| Reference project | This project |
|---|---|
for (i=0; i<TC6_MAX_INSTANCES; i++) loop finds first free slot; uses loop index i as TC6 instance number |
Uses DRV_LAN865X_DriverInfo *pDrvInst = (DRV_LAN865X_DriverInfo*)pGlobalTag and assigns pDrvInst->index as the TC6 instance number |
Reason: The PTP TX-Timestamp array drvTsCaptureStatus0[i] in
drv_lan865x_api.c is indexed by driver instance index. This change guarantees
TC6_t::instance == DRV_LAN865X_DriverInfo::index — a requirement for correct
timestamp attribution in multi-instance configurations.
#include <string.h>
#include "ptp_ts_ipc.h"
#include "ptp_fol_task.h"
#include "ptp_gm_task.h"
#define TCPIP_THIS_MODULE_ID TCPIP_MODULE_MANAGER
#include "library/tcpip/tcpip.h"
#include "library/tcpip/src/tcpip_packet.h"Registered with TCPIP_STACK_PacketHandlerRegister() on eth0. Intercepts frames
with EtherType 0x88F7 (PTP), acknowledges and returns true (consumed) so the
IP stack does not see them. No buffering — frame data is already captured by the
primary path (TC6_CB_OnRxEthernetPacket → g_ptp_raw_rx) at driver level
before this handler is called.
| State | Behaviour |
|---|---|
APP_STATE_INIT |
Prints build timestamp, sets state → APP_STATE_SERVICE_TASKS |
APP_STATE_SERVICE_TASKS |
Waits until TCPIP_STACK_NetIsUp() is true, then registers pktEth0Handler, sets state → APP_STATE_IDLE |
APP_STATE_IDLE |
First entry: calls PTP_FOL_Init(). Main loop: LAN register service, GM/FOL service every 1 ms, FOL frame delivery from driver path, LOFE reinit detection |
// GM: call every 1 ms
if (PTP_FOL_GetMode() == PTP_MASTER && (current_tick - last_gm_tick) >= ticks_per_ms) {
PTP_GM_Service();
}
// FOL: call every 1 ms
if (PTP_FOL_GetMode() == PTP_SLAVE && (current_tick - last_fol_tick) >= ticks_per_ms) {
PTP_FOL_Service();
}
// Deliver buffered PTP frame — filled by TC6_CB_OnRxEthernetPacket (driver level)
if (g_ptp_raw_rx.pending) {
g_ptp_raw_rx.pending = false;
if (PTP_FOL_GetMode() == PTP_SLAVE)
PTP_FOL_OnFrame(g_ptp_raw_rx.data, g_ptp_raw_rx.length, g_ptp_raw_rx.rxTimestamp);
}
// Re-run GM init after LAN865x LOFE recovery
bool lan865x_ready = DRV_LAN865X_IsReady(0u);
if (!lan865x_prev_ready && lan865x_ready && PTP_FOL_GetMode() == PTP_MASTER) {
PTP_GM_Init();
}
lan865x_prev_ready = lan865x_ready;Added APP_STATE_IDLE to the APP_STATES enumeration.
| Command | Description |
|---|---|
ptp_mode |
Show current PTP mode |
ptp_mode off |
Disable PTP |
ptp_mode master |
Enable Grandmaster mode |
ptp_mode master v |
Enable Grandmaster mode with per-FollowUp verbose line |
ptp_mode follower |
Enable Follower mode |
ptp_mode follower v |
Enable Follower mode with per-Sync verbose line |
ptp_status |
Show current mode, sync count, servo state, mean path delay |
ptp_time |
Show software PTP wallclock time and residual drift in ppb |
ptp_interval <ms> |
Set GM Sync interval in ms (default 125) |
ptp_offset |
Show follower clock offset (signed + absolute) in ns |
ptp_reset |
Reset follower servo to UNINIT |
ptp_trace on|off |
Enable/disable [TRACE] diagnostic messages on both GM and FOL |
ptp_dst |
Show current PTP destination MAC mode |
ptp_dst multicast|broadcast |
Set PTP destination MAC (default multicast) |
clk_set <ns> |
Force-set software wallclock to given nanosecond value, reset drift |
clk_get |
Read current software wallclock value in ns and drift in ppb |
lan_read <addr> |
Read LAN865x register at hex address |
lan_write <addr> <val> |
Write LAN865x register at hex address |
ptp_offset_reset |
Clear the HW-PTP offset ring buffer (used by ptp_offset_capture.py) |
ptp_offset_dump |
Dump all recorded HW-PTP offsets (one per line, <offset_ns> <sync_status>) |
loop_stats |
Show per-subsystem main-loop timing; loop_stats reset to clear |
sw_ntp_mode |
Show / set SW-NTP mode: sw_ntp_mode [off|master|follower <master_ip>] |
sw_ntp_poll <ms> |
Set follower poll interval in ms (default 1000, range 10..10000) |
sw_ntp_status |
Show SW-NTP mode, poll interval, sample count, timeouts, last offset |
sw_ntp_trace on|off |
Enable per-packet UART trace with all four SW timestamps (disturbs timing) |
sw_ntp_offset_reset |
Clear the SW-NTP offset ring buffer |
sw_ntp_offset_dump |
Dump all recorded SW-NTP offsets (one per line, <offset_ns> <valid>) |
tfuture_at <ns> |
Arm a firing event at absolute PTP_CLOCK nanosecond value |
tfuture_in <ms> |
Convenience: arm at now + <ms> |
tfuture_cancel |
Cancel a pending tfuture |
tfuture_status |
Show state, fires count, current drift_ppb, last target/actual/delta |
tfuture_reset |
Clear the tfuture ring buffer |
tfuture_dump |
Dump all recorded fires (one per line, <target_ns> <actual_ns> <delta>) |
tfuture_drift on|off |
Diagnostic: enable/disable drift correction in compute_target_tick |
ptp_gm_delay [<ns>] |
Diagnostic: add signed ns to GM anchor_wc at PTP_CLOCK_Update |
clk_set_drift [<ppb>] |
Diagnostic: manually force PTP_CLOCK drift_ppb |
cyclic_start [<period_us> [<anchor_ns>]] |
Start periodic GPIO toggle on PD10 at period_us (default 500 µs); shared anchor_ns makes GM and FOL edges phase-aligned. See §5.12. |
cyclic_stop |
Stop cyclic firing and drive PD10 low |
cyclic_status |
Show running flag, period, cycle count, miss count |
blink [<hz>|stop] |
Standalone PD10 rectangle generator: no arg → 1000 Hz, <hz> any positive integer, 0 or stop halts. Not PTP-synchronised. See §5.13. |
All periodic PTP prints use carriage-return-only (\r, no \n) to overwrite
the same terminal line continuously. State-transition messages are prefixed with
\r\n so they scroll normally without corrupting the overwriting line.
One overwriting line per FollowUp sent:
[GM] #498 t1=00:01:28.804334810
Format: [GM] #<seqId> t1=HH:MM:SS.nnnnnnnnn\r
The sec value (TSU counter in seconds) is decomposed to HH:MM:SS:
uint32_t h = sec / 3600u;
uint32_t m = (sec % 3600u) / 60u;
uint32_t s = sec % 60u;
SYS_CONSOLE_PRINT("[GM] #%u t1=%02lu:%02lu:%02lu.%09lu\r", ...);Activated with ptp_mode follower v. One overwriting line per received Sync:
[FOL] FINE t1=00:01:28.804334810 t2=00:01:28.804391620 off= +57 ns
Format: [FOL] <STATE> t1=HH:MM:SS.nnnnnnnnn t2=HH:MM:SS.nnnnnnnnn off=<±offset> ns delay=<delay> ns
State names (fixed-width 9 chars): UNINIT , MATCHFREQ, HARDSYNC ,
COARSE , FINE .
State transitions are printed with \r\n prefix so they do not overwrite the
running verbose line:
PTP_LOG("\r\nPTP COARSE offset=%d\r\n", (int)offset);
PTP_LOG("\r\nPTP FINE offset=%d\r\n", (int)offset);void PTP_FOL_SetVerbose(bool verbose); // ptp_fol_task.hCalled from ptp_mode_cmd() in app.c:
bool verbose = (argc >= 3) && (strcmp(argv[2], "v") == 0);
PTP_FOL_SetVerbose(verbose);The LAN865x hardware PTP clock (TSU counter) is accurate but can only be read
via an SPI register access — a blocking operation that takes several hundred
microseconds. ptp_clock.c creates a lightweight MCU-internal mirror of
that wallclock, queryable in nanoseconds with zero SPI traffic at query time.
This serves three purposes:
- Observability: query
clk_getvia UART from Python test scripts and compute the inter-board time difference directly, without touching the LAN865x over SPI. - Before/after comparison: by calling
clk_set 0on both boards simultaneously and then pollingclk_get, it is possible to measure the raw free-running crystal drift first (no PTP), and then repeat the same measurement with PTP active — demonstrating quantitatively what PTP synchronisation achieves (see §5.4). - Cross-board synchronisation: once the FOL servo has reached FINE state,
PTP_CLOCK_GetTime_ns()returns a network-wide consistent time on both boards. Application code can use this to schedule time-triggered actions, correlate events, or open coordinated measurement windows across the two MCUs without any additional synchronisation mechanism.
There are three distinct accuracy levels to distinguish:
The PTP servo controls the LAN8651 MAC_TSH/TSL registers directly. All four
timestamps t1–t4 originate from the hardware TSU, not the MCU.
- FINE threshold:
HARDSYNC_FINE_THRESHOLD = 500 ns— the servo enters FINE when the filtered offset falls below 500 ns; the FIR filter over the last 16 Sync samples holds the steady-state offset well below the threshold - Interpolation error between Syncs: 19 ppb residual drift × 125 ms = ~2.4 ns
- Dominant error source: path-delay asymmetry of the 10BASE-T1S segment (for a short point-to-point link < 1 m, typically < 10 ns)
Measured with ptp_offset_capture.py (§5.9), 60 s of FINE-state samples:
| Metric | Value |
|---|---|
|offset| mean |
46 ns |
| stdev | 24 ns |
| p95 | +2 ns (95 % of samples within ±50 ns) |
| worst case | 201 ns over 60 s |
The offset is computed from PHY-hardware timestamps directly in the firmware
(offset = (t2-t1) − mean_path_delay) and dumped via a ring buffer after
the measurement — no UART traffic during the run, no measurement artifacts.
See §5.9 for details.
PTP_CLOCK_GetTime_ns() returns t2_hardware + TC0_interpolation. The
wallclock component (t2) is the exact hardware timestamp; the anchor tick
sysTickAtRx is captured in the EIC EXTINT14 ISR on the nIRQ falling
edge (commit 5e289c8, see below), ~3–5 CPU cycles after the pin assertion.
Error sources:
- PHY-HW-Timestamp → nIRQ delay is ~80 µs for a Sync frame, constant per frame size and cancels out in the offset calculation.
- ISR-to-counter-read jitter: < 5 µs (vs. ~200 µs in the old polling design).
- TC0 rate error between Syncs: < a few ns at 125 ms intervals.
Achievable: < 5 µs inter-board accuracy of PTP_CLOCK_GetTime_ns().
Not directly measured in this project (would require an external reference
or a firmware-internal test that compares software-clock vs. hardware-clock
directly), but bounded above by the above error budget.
The clk_get-based tests (§5.1–5.8) see ~100 µs stdev. This is a
measurement infrastructure limit, not a board synchronicity number:
| Error source | Contribution |
|---|---|
| Python thread scheduling + USB-CDC polling | ~100 µs stdev (dominant) |
| Transport bursts (EDBG bridge congestion) | occasional ~9 ms outliers |
| Actual board synchronicity | hidden below the measurement floor |
The loop_stats instrumentation (§5.7) proves the firmware main loop never
blocks more than 209 µs, so the 9 ms outliers are entirely transport-path
artifacts — the firmware processes clk_get within sub-ms every time.
To see the real sub-µs / sub-100-ns synchronisation quality, use either:
ptp_offset_capture.py(§5.9) — firmware-internal PHY offset capture, or- A dual-channel oscilloscope on the two 1PPS outputs (external hardware).
The core idea is an anchor point: a pair (wallclock_ns, TC0_tick) captured
at the exact moment a hardware PTP timestamp arrives from the LAN865x.
anchor captured here
|
PTP wallclock: ------+---------------------------------------->
|<--- TC0 free-running since anchor --->|
|
PTP_CLOCK_GetTime_ns() |
= anchor_wc_ns
+ ticks_to_ns(now_tick - anchor_tick)
Between anchors the TC0 hardware timer (60 MHz, GCLK0/2) free-runs and provides sub-microsecond interpolation. The tick-to-nanosecond conversion is exact at this frequency — no rounding:
// 1 tick = 1e9 / 60e6 ns = 50/3 ns (exact integer ratio)
static uint64_t ticks_to_ns(uint64_t ticks)
{
return (ticks / 3ULL) * 50ULL + ((ticks % 3ULL) * 50ULL) / 3ULL;
}Anchors arrive every ~125 ms (one per Sync frame). The maximum interpolation error within a 125 ms window for a crystal running at ±500 ppm is at most ±62 µs — well below the re-anchoring residual noise of ±65–130 µs.
| Board role | Called from | Hardware event |
|---|---|---|
| Follower | ptp_fol_task.c via PTP_CLOCK_Update() |
RX timestamp from RTSA footer (stripped by TC6 driver) on every Sync/FollowUp |
| Grandmaster | ptp_gm_task.c via PTP_CLOCK_Update() |
TX timestamp from TTSCAH/AL registers after Sync frame sent (TTSCAA bit) |
In both cases PTP_CLOCK_Update(wallclock_ns, sys_tick) stores the pair:
void PTP_CLOCK_Update(uint64_t wallclock_ns, uint64_t sys_tick)
{
s_anchor_wc_ns = wallclock_ns; // hardware PTP timestamp (ns)
s_anchor_tick = sys_tick; // TC0 tick at that exact moment
s_valid = true;
}ptp_sync_before_after_test.py measures the inter-board time difference before
and after enabling PTP. Each phase runs for 60 s with 500 ms sample interval;
a linear regression separates the systematic frequency drift from residual
noise.
Test run — 2026-04-16, boards SAME54+LAN8650, COM8 (GM) / COM10 (FOL):
| Metric | Phase 0 — Free-run | Phase 1 — PTP active |
|---|---|---|
| Slope (ppb) | −32 450 | +208 |
| Slope (ppm) | −32.45 | +0.21 |
| Residual stdev | 266 µs | 107 µs |
| Drift FOL | 0 ppb | −19 ppb (±31 ppb stdev) |
| Samples valid | 116/118 | 116/118 |
| FINE convergence | — | 2.8 s (HardSync@0.4 s, MATCHFREQ@2.3 s) |
| Slope reduction | — | 99.4 % |
Without PTP the FOL crystal ran 32.45 ppm slower than the GM — the clocks would diverge by ~2 ms per minute, ~117 ms per hour. With PTP active the residual slope collapses to 0.21 ppm (99.4 % reduction). The remaining slope is the servo's residual frequency error after the one-time TISUBN crystal calibration; it resets toward zero over several Sync frames.
The residual stdev of 107 µs reflects the measurement floor of reading
clk_get from two independent COM ports via Python on Windows. The two UART
reads are not strictly simultaneous — the elapsed PC time between them (USB
polling, OS scheduling) contributes directly to the measured difference and
sets the lower bound for what the test can resolve.
Two samples spike to ≈+9 ms (+9078 µs at t = 4.1 s, +9134 µs at t = 53.2 s). Both are removed by the outlier filter (2/118 = 1.7 %).
These are PC-side measurement artifacts, not hardware clock errors. The
test reads clk_get from COM8 (GM) and COM10 (FOL) in sequence using two
separate USB-to-UART adapters. Windows USB polling runs at 8 ms intervals by
default. On rare occasions the OS thread is preempted or a USB frame is delayed
between the two reads, introducing a gap of ~8–16 ms that appears directly as
a spike in the measured inter-board difference. The outlier filter (IQR-based)
reliably removes these before the regression.
After clk_set 0 on both boards simultaneously (anchor set once via
PTP_CLOCK_ForceSet(0), never updated again), TC0 runs freely on each board's
independent crystal:
[Board GM] clk_get: 29 786 000 000 ns drift=0ppb
[Board FOL] clk_get: 29 784 031 700 ns drift=0ppb
difference: ≈−1968 ns after ~60 s (−32.45 ppm × 60 s)
The difference grows linearly at the crystal frequency error between the two
boards. Free-run slope depends on the specific board pair; values from −262 ppm
to −32 ppm have been measured across different runs. This is the Phase 0
baseline captured by ptp_sync_before_after_test.py.
Once the FOL servo reaches FINE state, PTP_CLOCK_Update() is called every
~125 ms with the hardware PTP timestamp. TC0 interpolates in between.
Both boards now track the same shared time:
[Board GM] clk_get: 3141592653 ns drift=0ppb
[Board FOL] clk_get: 3141592610 ns drift=-19ppb
difference: ≈−43 ns (within FINE threshold ±500 ns)
drift_ppb (non-zero on FOL only) shows the residual frequency error the
PTP servo is still trimming after the one-time TISUBN crystal calibration.
Measured: −19 ppb mean, ±31 ppb stdev in Phase 1.
The FOL servo writes the crystal-calibrated value to MAC_TISUBN once at
UNINIT→MATCHFREQ. This corrects the LAN865x TSU counter frequency to match the
GM crystal. From MATCHFREQ onward the hardware-timestamp-based anchors are
frequency-matched, and the remaining interpolation error is only UART
scheduling jitter (107 µs stdev measured in Phase 1 of the before/after
test).
After the TISUBN correction, drift_ppb reports the remaining frequency error
observed by the PTP servo (in parts per billion):
int32_t PTP_CLOCK_GetDriftPPB(void); // read — used by clk_get CLI
void PTP_CLOCK_SetDriftPPB(int32_t); // written by ptp_fol_task.c after each FIR updateUpdated in ptp_fol_task.c at every Sync frame in COARSE or FINE:
PTP_CLOCK_SetDriftPPB((int32_t)((rateRatioFIR - 1.0) * 1e9));Observed values: −19 ppb mean, ±31 ppb stdev residual after TISUBN correction.
Resets to 0 on clk_set 0 (PTP_CLOCK_ForceSet()).
The drift estimate inside ptp_clock.c is filtered with an adaptive
single-pole IIR: the effective filter window grows from 1 (sample 1) up to
the configured ceiling s_drift_iir_n (default 128) as samples accumulate.
This breaks the usual single-pole settle-vs-jitter trade-off: warm-up converges
in sub-second time, while the steady-state jitter floor stays at the 1/√N_max
level (~7 µs MAD on cyclic_fire cross-board edges, measured 2026-04-23).
Two CLI commands tune the filter at runtime:
| Command | Effect |
|---|---|
drift_iir_n [<8..4096>] |
Get/set the steady-state ceiling N_max (default 128). Larger N → lower jitter floor, longer ramp to that floor. Warm-up speed is unaffected. |
drift_iir_reset |
Re-arm the warm-up ramp (sample counter → 0). Used by test scripts to make settle-time measurements reproducible. |
Full design rationale, measurement methodology, fixed-vs-adaptive comparison, and tuning guidance: documentation/ptp/drift_filter.md
Once the FOL servo reaches FINE state, both boards share a common timebase
accurate to ±500 ns (current HARDSYNC_FINE_THRESHOLD). Application code
can use PTP_CLOCK_GetTime_ns() for any purpose that requires events on
different nodes to be correlated in time:
Time-triggered actions — both boards execute an action at the same absolute PTP timestamp:
uint64_t fire_at_ns = 5000000000ULL; // T = 5 s after PTP epoch
while (PTP_CLOCK_GetTime_ns() < fire_at_ns) { /* spin */ }
GPIO_PA01_Toggle(); // fires on both boards within ±500 ns of each otherCorrelated event logging — timestamps from different boards are directly comparable:
PTP_LOG("[EVENT] ts=%llu ns\r\n",
(unsigned long long)PTP_CLOCK_GetTime_ns());Coordinated measurement windows — start a measurement at the next full second boundary, guaranteed to be the same second on both boards:
uint64_t now = PTP_CLOCK_GetTime_ns();
uint64_t next = ((now / 1000000000ULL) + 1ULL) * 1000000000ULL;
while (PTP_CLOCK_GetTime_ns() < next) { /* wait */ }
start_measurement();Important:
PTP_CLOCK_IsValid()must returntruebefore callingPTP_CLOCK_GetTime_ns()for synchronized purposes. On the FOL this is guaranteed only after the first FollowUp frame has been processed (i.e. after the servo leaves UNINIT).PTP_CLOCK_GetTime_ns()is not interrupt-safe — do not call it from an ISR.
| Command | Description |
|---|---|
ptp_time |
Print current wallclock as HH:MM:SS.ns and residual drift: ptp_time: HH:MM:SS.nnnnnnnnn drift=<±ppb>ppb |
clk_get |
Print raw wallclock in nanoseconds and drift: clk_get: <ns> drift=<±ppb>ppb |
clk_set 0 |
Zero the wallclock (independent timer baseline for before/after tests) |
Before commit 5e289c8, the LAN865x nIRQ line (PC14) was polled inside
DRV_LAN865X_Tasks(). When the pin was found asserted, an SPI transfer read
the pending frame(s) from the FIFO. sysTickAtRx was captured inside
TC6_CB_OnRxEthernetPacket() at the end of that SPI transfer:
LAN8651 SFD received → TSU latches hw_rx_timestamp
↓ (frame bytes arrive, ~80 µs for 100-byte frame @ 10 Mbps)
nIRQ PC14 asserts LOW
↓ ← polling jitter (0 … several ms, variable)
DRV_LAN865X_Tasks() → SYS_PORT_PinRead(PC14) == LOW
↓ SPI transfer (~100–300 µs, deterministic)
TC6_CB_OnRxEthernetPacket() → sysTickAtRx = SYS_TIME_Counter64Get()
↓
PTP_CLOCK_Update(hw_rx_timestamp, sysTickAtRx)
The polling jitter between nIRQ assertion and SPI completion landed directly
in sysTickAtRx and therefore in every anchor pair — the anchor tick could be
off by a variable amount each Sync cycle.
PC14 is routed through the ATSAME54 EIC peripheral (EXTINT[14], falling edge).
A minimal ISR captures the TC0 tick at the moment nIRQ asserts — before
any SPI activity starts:
/* EIC EXTINT14 ISR — fires on nIRQ falling edge, ISR latency ~3-5 CPU cycles */
void EIC_EXTINT_14_Handler(void)
{
s_nirq_tick = SYS_TIME_Counter64Get(); /* capture tick immediately */
s_nirq_pending = true;
EIC_REGS->EIC_INTFLAG = EIC_INTFLAG_EXTINT14_Msk;
}DRV_LAN865X_Tasks() now polls s_nirq_pending instead of the raw pin, and
TC6_CB_OnRxEthernetPacket() uses the ISR-captured s_nirq_tick for
sysTickAtRx instead of reading the counter again at SPI completion.
The anchor pair becomes (hw_rx_timestamp, nirq_tick). The fixed delay
between hardware timestamp and nIRQ assertion (≈80 µs for a 100-byte Sync
frame) is constant per frame and cancels out in the servo's offset
calculation t2_local − t1_GM.
Two additional details of the implementation:
_InitNIrqEIC()configures EIC clock, sets EXTINT14 to SENSE=FALL, and enables the NVIC.PORT_PINCFG[14]was changed from0x6to0x7so that PC14 is routed to the EIC peripheral function (alongside the GPIO read, which still works).- If
nIRQis still low afterTC6_Service()returns (a second falling edge that arrived during the service call),s_nirq_pendingis re-armed so the next iteration handles it.
For a typical PTPv2 Sync frame (~64 bytes Ethernet, 10 Mbps line rate):
T0 : SFD detected on MDI → TSU latches PHY-HW-Timestamp (t2, sub-ns precise)
│
│ Frame transmission on the wire:
│ 64 B payload × 8 bits / 10 Mbps = 51 µs
│ + preamble + SFD (8 B) = 6 µs
│ + FCS (4 B) = 3 µs
│ ≈ 60 µs
│
T0 + 60 µs : frame complete; LAN8651 packs into internal FIFO
│ PHY-internal delay until nIRQ ≈ 5–20 µs
│
T0 + 80 µs : nIRQ PC14 falls
│ Cortex-M4 @ 120 MHz IRQ entry ≈ 100 ns
│ + prologue to function body ≈ 50 ns
│
T0 + 80 µs + 150 ns : enter EIC_EXTINT_14_Handler
│ SYS_TIME_Counter64Get()
│ (SYS_INT_Disable + HW-read + Restore) ≈ 200–500 ns
│
T0 + 80 µs + 500 ns : s_nirq_tick stored
Total PHY-to-MCU-tick delay ≈ 80 µs, dominated by the wire transmission
time. This is a deterministic offset per frame size — it cancels out in
the PTP math (offset = t2 − t1 − delay) because both GM and FOL experience
the same constant-per-frame delay. Only the jitter around the 80 µs mean
matters, and the ISR path keeps that below ~5 µs.
The improvement reduced sysTickAtRx jitter from ~200 µs to <5 µs, but this
is not visible in the UART/Python test suite. The measurement floor of
ptp_drift_compensate_test.py is ~100–200 µs (Windows USB-CDC polling), and
occasional 9 ms outliers appear as well — these have been traced to
UART/USB-CDC transport jitter, not PTP. See §5.7 below and the loop_stats
instrumentation which proves the firmware main-loop never blocks for >209 µs.
Further accuracy gains would become measurable with an oscilloscope comparing the two boards' 1PPS outputs.
File: src/config/default/initialization.c
Before calling TCPIP_STACK_Init(), the last three bytes of the Ethernet MAC
address are randomised using the ATSAME54 hardware TRNG peripheral.
Changes:
#include <stdio.h>added.s_macAddrStr0[18]buffer declared at file scope.APP_RandomizeMacLastBytes()function added:- Enables MCLK for TRNG (
MCLK_APBCMASK_TRNG_Msk). - Enables TRNG (
TRNG_CTRLA_ENABLE_Msk). - Polls
TRNG_INTFLAG_DATARDY_Msk, readsTRNG_DATA. - Formats result as
"00:04:25:XX:XX:XX"intos_macAddrStr0.
- Enables MCLK for TRNG (
- Called immediately before
TCPIP_STACK_Init(). TCPIP_HOSTS_CONFIGURATION[0].macAddrfield set tos_macAddrStr0.
File: src/app.c
Two CLI commands added to the Test command group for run-time LAN865x SPI
register access without a debugger.
| Command | Description |
|---|---|
Test lan_read <addr_hex> |
Read a LAN865x register and print the result |
Test lan_write <addr_hex> <value_hex> |
Write a LAN865x register |
Implementation details:
- Non-blocking state machine:
app_lan_state_t(IDLE / WAIT_READ / WAIT_WRITE). - Callbacks
lan_read_callback()/lan_write_callback()set volatile flags. - 200 ms timeout (
APP_LAN_TIMEOUT_MS) viaSYS_TIME_Counter64Get(). - Commands call
DRV_LAN865X_ReadRegister()/DRV_LAN865X_WriteRegister()on driver instance 0. Command_Init()registers the group viaSYS_CMD_ADDGRP(), called fromAPP_Initialize().
The ATSAME54P20A firmware exposes a Harmony CLI over UART: a simple
line-based command shell accessible at 115200 baud via the standard SYS_CMD
module. All test scripts remotely drive this CLI over pyserial, without any
debugger connection or custom firmware protocol.
The general pattern used by every script:
Host PC (Python) Board (ATSAME54 firmware)
─────────────────────────────────────────────────────────────────
ser.write(b"ptp_mode follower\r\n") → Harmony CLI processes command
SYS_CONSOLE_PRINT("[PTP-FOL] ...")
response = ser.read_until(timeout) ← UART TX: status / event strings
Parsing is done by regex on the UART output — no binary protocol, no
custom framing. Each script opens two serial.Serial instances (one per
board), sends commands and reads responses in parallel Python threads where
necessary (e.g. simultaneous clk_set 0 for baseline synchronisation).
Key constraint on Windows: never read the same pyserial.Serial port
from two threads simultaneously — this corrupts the Win32 overlapped I/O
handles and causes heap corruption. Every script either uses a single reader
thread per port, or serialises port access at the call site (see §5.2 for the
specific fix applied to ptp_onoff_test.py).
Prerequisites: pip install pyserial
Board port assignment (default): Board 1 = COM8, Board 2 = COM10
(as configured by setup_flasher.py).
Run the tests in the order shown — §5.1 is the sanity baseline, §5.2–§5.3 verify specific fix correctness, §5.4–§5.5 demonstrate end-to-end accuracy.
Validates that the TC0-based software clock (PTP_CLOCK) operates correctly
without PTP. Measures the raw crystal frequency difference between the two
boards and verifies that the TC0 tick-to-nanosecond interpolation is internally
consistent. Run this first as a sanity check before any PTP test.
| What it proves | What it does NOT prove |
|---|---|
| TC0 tick-to-ns conversion correct on both boards | Correct PTP anchor capture |
PTP_CLOCK_GetTime_ns() interpolation consistent |
PTP Ethernet timestamping (RTSA / TTSCAL) |
| UART serialisation latency correction works | |
| Crystal frequency ratio measured accurately |
| Step | Description |
|---|---|
0 — Simultaneous clk_set 0 |
Zeroes both clocks in parallel threads; measures thread launch skew. |
| 1 — Settle | 2 s pause. |
| 2 — Collect 100 paired samples | 100 swap-symmetrised clk_get pairs at 100 ms intervals. Linear regression removes crystal trend; residuals must be < 500 µs. |
The growing mean offset (linear drift) is expected — two free-running crystal oscillators always drift apart at a constant rate. The PASS check only verifies the residuals after removing that linear trend.
python hw_timer_sync_test.py --a-port COM8 --b-port COM10Optional: --n <samples> (default 100), --pause-ms (default 100),
--threshold-us (default 500), --settle-s (default 2), --log-file <path>.
| Run | Date/time | Slope (ppm) | Intercept (µs) | Res. stdev (µs) | Result |
|---|---|---|---|---|---|
| 1 | 2026-04-09 17:56 | −321.2 | −687 | 91 | PASS |
| 2 | 2026-04-09 17:57 | −266.0 | −424 | 134 | PASS |
| 3 | 2026-04-10 09:08 | −0.0¹ | −22 | 145 | PASS |
| 4 | 2026-04-10 09:08 | −392.4 | −453 | 192 | PASS |
| 5 | 2026-04-10 09:12 | +5.5 | −66 | 62 | PASS |
| 6 | 2026-04-10 09:19 | −139.9 | −691 | 149 | PASS |
¹ Run 3 immediately after a PTP session; TISUBN retained its correction → apparent slope ≈ 0. Residual stdev still 145 µs < 500 µs.
All 6 runs: PASS (3/3 steps each)
Automated resilience test: starts PTP, measures baseline offset, stops GM, observes drift, restarts GM, verifies re-convergence and post-restart accuracy.
python ptp_onoff_test.py --gm-port COM10 --fol-port COM8| Phase | Description |
|---|---|
| Steps 0–3 | Reset boards, set IP addresses, bidirectional ping, start PTP, wait for FOL FINE |
| Phase A | Baseline: 10 offset samples while GM running |
| Phase B | Blackout: stop GM, monitor FOL offset for 5 s |
| Phase C | Re-convergence: restart GM, wait for FINE |
| Phase D | Post-restart: 10 offset samples, verify ±100 ns |
| Pattern | Firmware string matched |
|---|---|
RE_HARD_SYNC |
"Hard sync completed" |
RE_MATCHFREQ |
"UNINIT->MATCHFREQ" |
RE_COARSE |
"PTP COARSE" |
RE_FINE |
"PTP FINE" |
Rule: Never read the same pyserial.Serial port from two threads
simultaneously on Windows — corrupts overlapped I/O handles → access violation.
# 1. FOL command — main thread owns fol_ser exclusively
resp = send_command(self.fol_ser, "ptp_mode follower", ...)
time.sleep(0.5)
# 2. Start convergence thread (sole reader of fol_ser from here on)
self.fol_ser.reset_input_buffer()
self._start_convergence_thread()
# 3. GM command — uses gm_ser only, safe to run concurrently
self.gm_ser.write(b"ptp_mode master\r\n")| Phase | Result |
|---|---|
| Initial convergence | FINE in 2.7 s (HARD_SYNC @ 0.4 s, MATCHFREQ @ 2.3 s) |
| Phase A baseline (n=10) | mean = +36.5 ns, stdev = 16.5 ns |
| Phase B blackout (5 s) | Offset frozen at +78 ns, drift = 0 ns — local clock holds |
| Phase C re-convergence | FINE in 0.8 s — saved TI/TISUBN reused, MATCHFREQ skipped |
| Phase D post-restart (n=10) | mean = +43.9 ns, stdev = 14.9 ns, 10/10 within ±100 ns |
Overall: PASS (5/5)
Verifies correct convergence after a GM↔FOL role swap. Directly validates Bug Fix §2.2.
python ptp_role_swap_test.py --board1-port COM8 --board2-port COM10Optional: --board1-ip, --board2-ip, --convergence-timeout (default 30 s),
--verbose.
| Phase | Board 1 | Board 2 |
|---|---|---|
| Phase 1 | Follower | Grandmaster — wait for FINE, collect 10 offset samples |
| — | ptp_mode off on both, pause 5 s |
|
| Phase 2 | Grandmaster | Follower — wait for FINE, collect 10 offset samples |
- Both phases reach FINE within
--convergence-timeout - Phase 2 mean offset within ±500 ns, stdev < 100 ns, ≥ 9/10 samples within ±500 ns
Before fix:
| Metric | Phase 1 | Phase 2 |
|---|---|---|
| FINE reached | 2.7 s | never (timeout at 30 s) |
| mean offset | +52 ns | −3 138 186 ns (stuck) |
After fix (2 confirmed runs):
| Run | Ph.1 FINE | Ph.1 mean | Ph.2 FINE | Ph.2 mean | Ph.2 stdev | ≤±500 ns |
|---|---|---|---|---|---|---|
| 1 | 2.7 s | +50.8 ns | 2.7 s | −4.0 ns | 12.8 ns | 10/10 |
| 2 | 3.1 s | +57.9 ns | 2.9 s | −9.2 ns | 21.2 ns | 10/10 |
Overall: PASS (6/6 checks, both runs)
Demonstrates PTP synchronisation quantitatively: free-running crystal drift measured first, then again with PTP active — both in a single automated run.
| Phase | Description |
|---|---|
| Phase 0 — Free Running | Both clocks zeroed simultaneously, clk_get pairs collected for 60 s. Linear regression shows raw crystal drift. |
| PTP Setup | IP config, start Follower + Grandmaster, wait for FINE. |
| Phase 1 — PTP Active | Clocks re-zeroed, clk_get pairs collected for 60 s with PTP running. |
| Comparison | Side-by-side table: slope, residual stdev, drift reduction %. |
Measurement uses swap symmetry: alternating samples query GM first then FOL, the next queries FOL first then GM. The send-time skew between the two parallel threads is subtracted, eliminating systematic host PC scheduler bias.
python ptp_sync_before_after_test.py --gm-port COM8 --fol-port COM10| Criterion | Default threshold |
|---|---|
| ` | slope_ptp |
| Residual stdev | < 500 µs |
| Run | Free-run (ppm) | PTP slope (ppm) | PTP stdev (µs) | drift stdev (ppb) | FINE (s) |
|---|---|---|---|---|---|
| 1 | +179.3 | −0.243 | 68 | 0¹ | 2.8 |
| 2 | +209.0 | −0.183 | 66 | 0¹ | 2.9 |
| 3 | −75.4 | +0.146 | 129 | 0¹ | 2.7 |
| 4 | −518.5 | +0.666 | 92 | 6 | 2.7 |
| 5 | −229.3 | +0.641 | 121 | 10 | 2.7 |
| 6–10² | −532…−46 | −0.43…+0.75 | 60…78 | 6–10 | 2.7 |
¹ Runs 1–3 predated the drift_ppb firmware update (value was always 0).
² Runs 6–10 from the reproducibility test (§5.5).
All 10 runs: PASS
Runs ptp_sync_before_after_test.py N times (default 5) and aggregates all
results in a single summary table. Purpose: automated verification of
reproducibility.
python ptp_reproducibility_test.py --gm-port COM8 --fol-port COM10
python ptp_reproducibility_test.py --gm-port COM8 --fol-port COM10 --runs 3All arguments (--free-run-s, --ptp-s, --slope-threshold-ppm, etc.) are
forwarded to the sub-test.
Each run generates a unique log filename
(ptp_sync_before_after_test_<YYYYMMDD_HHMMSS>.log) before starting the
sub-test and passes it via --log-file — no directory search, no ambiguity.
| Column | Description |
|---|---|
Free(ppm) |
Free-run crystal drift (Phase 0 slope) |
FreeStd |
Free-run residual stdev |
FINE(s) |
PTP convergence time to FINE state |
PTP(ppm) |
PTP-active slope (Phase 1) |
PTPStd |
PTP-active residual stdev |
dFOLstd |
Follower drift_ppb stdev during Phase 1 |
Reduc% |
Slope reduction by PTP in percent |
Dur(s) |
Single-run wall-clock duration |
Result |
PASS / FAIL |
5 consecutive runs, 60 s free-run + 60 s PTP each, total 12 minutes:
| Run | Free (ppm) | PTP (ppm) | PTP stdev | FINE (s) | Reduc. |
|---|---|---|---|---|---|
| 1 | −532.5 | +0.502 | 60 µs | 2.7 s | 99.9% |
| 2 | −517.6 | +0.198 | 61 µs | 2.7 s | 100.0% |
| 3 | −293.3 | +0.745 | 61 µs | 2.7 s | 99.7% |
| 4 | −365.7 | +0.306 | 78 µs | 2.7 s | 99.9% |
| 5 | −45.7 | −0.429 | 62 µs | 2.7 s | 99.1% |
PTP slope: mean = +0.26 ppm · stdev = ±0.44 ppm
PTP stdev: mean = 64 µs · stdev = ±7.6 µs
FINE time: 2.7 s in all 5 runs
Overall: PASS (5/5)
End-to-end regression test that verifies full PTP convergence and the complete Delay_Req/Delay_Resp exchange with hardware timestamp capture. Unlike §5.2–5.5 (which measure offset quality over time), this test focuses on protocol correctness and is the primary verification tool after any change to the PTP frame-building or timestamp-capture code.
ptp_trace onis activated immediately after PTP start — before the first Sync frame is even sent — so no trace event is ever missed.- Both serial ports are read by permanent background threads from the start; there is no read race between convergence polling and trace capture.
- The test does not abort if FINE is not reached — trace analysis and
assertions run regardless, followed by a detailed
STUCK-STATE DIAGNOSEsection. - Convergence timeout is 60 s (vs. 30 s in earlier scripts).
python ptp_trace_debug_test.py --gm-port COM10 --fol-port COM8
python ptp_trace_debug_test.py --gm-port COM10 --fol-port COM8 ^
--convergence-timeout 90 --trace-time 20| Step | Action | Pass condition |
|---|---|---|
| 0 | Reset both boards, wait 8 s | Boot completes |
| 1 | setip eth0 on GM + FOL |
IP set confirmed |
| 2 | ping bidirectional |
Ping: done. on both sides |
| 3 | ptp_mode follower → ptp_trace on (FOL, immediately) |
Trace enabled before first Sync |
| 3 | ptp_mode master → ptp_trace on (GM, immediately) |
Trace enabled |
| 3 | Poll FOL for PTP FINE (≤ 60 s) |
FINE reached; milestones logged |
| 4 | Collect 10 s additional trace | All Delay exchanges captured |
| 5 | ptp_trace off, ptp_mode off |
Clean shutdown |
| Assertion | What it verifies |
|---|---|
| A | FOL sent at least one DELAY_REQ_SENT |
| B | GM received at least one GM_DELAY_REQ_RECEIVED |
| C | GM sent at least one GM_DELAY_RESP_SENT |
| D | FOL received at least one DELAY_RESP_RECEIVED |
| E | At least one DELAY_CALC shows non-zero, plausible delay |
| F | GM_DELAY_RESP_SKIPPED_TX_BUSY count ≤ limit (default 0) |
| G | Last valid delay in range 0 < delay < 10 ms |
| H | At least one DELAY_CALC shows hw=1 (t3 from LAN865x TTSCA) |
| I | Zero DELAY_RESP_WRONG_SEQ events (IEEE 1588 §11.3.3 seq-ID check) |
[PASS] Step 3: PTP Start + ptp_trace ON (immediately) + Convergence — FINE@2.7s
[PASS] A: FOL DELAY_REQ_SENT count=45
[PASS] B: GM GM_DELAY_REQ_RECEIVED count=45
[PASS] C: GM GM_DELAY_RESP_SENT count=45
[PASS] D: FOL DELAY_RESP_RECEIVED received=45
[PASS] E: FOL DELAY_CALC non-zero delay last_valid_delay=3788 ns
[PASS] F: GM TX-busy skips <= limit skips=0 limit=0
[PASS] G: Delay in plausible range 3788 ns
[PASS] H: FOL t3 HW-Capture (hw=1) hw_captures=45/45
[PASS] I: No DELAY_RESP_WRONG_SEQ no WRONG_SEQ — seq-ID check correct
OVERALL: PASS
Validates that PTP synchronisation actively compensates the crystal frequency
difference between GM and FOL, measured through the TC0-based software clock
(clk_get CLI command). Unlike §5.1–5.6 which either test convergence or
protocol correctness, this test runs the servo for 120 s and checks that the
residual diff(t) = FOL_clk − GM_clk stays bounded.
| Step | Action | Pass condition |
|---|---|---|
| 0 | Reset both boards | boot completes |
| 1 | Set IPs via setip |
both IPs confirmed |
| 2 | Ping bidirectional | both sides reply |
| 3 | ptp_mode master/follower, wait for FINE |
FINE in ≤60 s |
| 4 | Enable ptp_trace + start SerialMux background readers |
mux routes clk_get: → clk queue, rest → trace queue |
| 5 | clk_set 0 on BOTH boards in parallel (thread rendezvous) |
thread send skew < 1 ms |
| 6 | Collect 120 s of paired clk_get samples at 500 ms intervals |
slope, stdev, outlier count |
| 7 | Query firmware loop_stats — max/avg per-subsystem main-loop time |
max TOTAL < 1 ms |
| Metric | Threshold | Typical |
|---|---|---|
|slope| of diff(t) |
2.0 ppm | 0.01–0.35 ppm |
| residual stdev (after detrending) | 500 µs | 100–200 µs |
Occasional diff samples show +9 ms spikes:
- With
ptp_trace on: ~1 outlier per 30 s - Without
ptp_trace: ~1 outlier per 4 min
Two firmware instruments rule out a PTP software fault:
loop_statsCLI command (added inloop_stats.c) records max/avg time spent in each subsystem of the Harmony super-loop (SYS_CMD_Tasks,TCPIP_STACK_Task,ptp_log_flush,APP_Tasks, andTOTALiteration). A 120 s test with trace enabled reports:max_TOTAL = 209 µs,avg_TOTAL = 21 µsacross 5.3 million iterations — the main loop never blocks for more than 0.21 ms.ptp_clock.creads the TC0 hardware counter (60 MHz) directly at the momentclk_get_cmdruns. It is anchored by each PTP Sync (every 125 ms) and linearly interpolates between Syncs via TC0 ticks. There is no code path where the clock reading could be 9 ms in the future.
Conclusion: the 9 ms is the difference between the moments the two CPUs
process their respective clk_get commands — caused by UART/USB-CDC
transport jitter at the EDBG bridge chip, not by any PTP error.
loop_stats.c/.h— per-subsystem timing viaSYS_TIME_Counter64Get()pairs wrapped around each super-loop call.loop_statsCLI command (loop_stats/loop_stats reset).- Async Delay_Req timeout: moved from
sendDelayReq()(called every ~125 ms on FollowUp) intoPTP_FOL_Service()(called every 1 ms), so the 500 ms timeout is detected with 1 ms granularity instead of up to 125 ms late. - Rate-limited
ptp_log_flush(): drains at mostPTP_LOG_FLUSH_PER_TICK = 2messages per super-loop iteration so a burst of queued trace lines can't starveSYS_CMD_Tasks.
python ptp_drift_compensate_test.py --gm-port COM8 --fol-port COM10
python ptp_drift_compensate_test.py --gm-port COM8 --fol-port COM10 --traceBy default ptp_trace is OFF (trace amplifies UART-induced outliers —
see §5.7 key finding). The SerialMux architecture is always used so
clk_get polling is thread-safe. Pass --trace to send ptp_trace on to
the firmware during the collection window when you actually need the PTP
event stream for diagnosis.
Like §5.4 (ptp_sync_before_after_test.py) this test demonstrates PTP
synchronisation in a single run by collecting paired clk_get samples in
two phases and printing a before/after comparison. The difference is the
mux-based architecture borrowed from §5.7:
SerialMuxbackground reader thread per port (safeclk_geteven with concurrentptp_trace)- Parallel
clk_set 0via thread-rendezvous → sub-ms skew at the start of each phase (instead of ~31 ms sequential skew in the older test) - Prompt/echo filter → clean
[PTP][…]lines without>clk_getnoise loop_statsquery at the end of each phase to prove the firmware main-loop is never the bottleneck
| Step | Action |
|---|---|
| 0 | Reset both boards, wait 8 s |
| Phase 0 | Start mux, parallel clk_set 0, collect --free-run-s s of samples. Query loop_stats, stop mux. |
| 1 | setip eth0 on GM + FOL |
| 2 | ping bidirectional |
| 3 | ptp_mode master/follower, wait for FOL FINE |
| Phase 1 | (optional ptp_trace on), start mux, settle, parallel clk_set 0, collect --ptp-s s. Query loop_stats, stop mux, ptp_trace off. |
| Compare | Side-by-side table: free vs. PTP (slope ppb/ppm, residual stdev, drift_fol) + reduction %. |
| Metric | Threshold |
|---|---|
|slope| of diff(t) |
2.0 ppm |
| residual stdev (after detrending) | 500 µs |
| overall (comparison) | slope reduction > 50 % OR |slope_ptp| below threshold |
Metric Free-run (no PTP) PTP active
Slope (ppb) +3900 +913
Slope (ppm) +3.9001 +0.9131
Residual stdev (us) 590.049 101.059
Slope reduction by PTP : 76.6 %
...
loop_stats: max TOTAL = 142 us over 2.7M iterations
Longer --ptp-s 120 drives the IIR to full convergence and typically
pushes the residual slope below 0.1 ppm (reduction ≥ 97 %).
python ptp_sync_before_after_mux_test.py --gm-port COM8 --fol-port COM10
python ptp_sync_before_after_mux_test.py --gm-port COM8 --fol-port COM10 ^
--free-run-s 120 --ptp-s 120
python ptp_sync_before_after_mux_test.py --gm-port COM8 --fol-port COM10 --traceThe test imports its helpers directly from ptp_drift_compensate_test.py
(SerialMux, TraceCollector, collect_clk_get_samples, regression helpers)
— both test scripts must live in the same directory.
Tests §5.1–5.8 all measure via the clk_get CLI command, which is limited
by UART/USB-CDC transport jitter (~100 µs floor, occasional ~9 ms outliers).
This test bypasses that limit entirely: the Follower firmware records the
raw PHY-hardware PTP offset offset = (t2-t1) - mean_path_delay into a
ring buffer every Sync cycle (~125 ms), and the CLI dumps all samples
after the measurement is done. No UART traffic during the measurement
→ no measurement distortion.
Level 1: PHY Hardware Timestamps (this test)
t1, t2, t3, t4 — latched by LAN8651 TSU at SFD, sub-ns resolution
offset = (t2 - t1) - ((t2-t1) + (t4-t3)) / 2
Measured sync quality: |offset| mean = ~46 ns, stdev = ~24 ns
│
▼ anchor_wc_ns = t2, anchor_tick = sysTickAtRx (EIC ISR)
Level 2: Software Wallclock Timer (PTP_CLOCK_GetTime_ns)
Inherits PHY precision as anchor, plus ~5 µs sysTickAtRx jitter,
plus TC0 rate error between Syncs. Not directly measured here.
│
▼ clk_get CLI command via UART / USB-CDC
Level 3: CLI Measurement (§5.7, §5.8)
Inherits Level 2, plus ~100 µs USB-CDC jitter, plus rare transport
bursts (9 ms outliers). Sync quality APPEARS ~100 µs even when the
underlying PTP is ~50 ns.
This test pins down Level 1 — the actual PTP protocol accuracy.
ptp_offset_trace.c/.h— ring buffer of 1024 × (int32 offset_ns + uint8 sync_status), 5 KB RAM. Populated fromprocessFollowUp()right after the offset calculation.- CLI commands added to
app.c:ptp_offset_reset— clear the ring buffer before a measurement.ptp_offset_dump— print all samples, one per line, as<offset_ns> <sync_status>. Rate-limited (4 lines per batch, 20 ms pause) soSYS_CONSOLE_PRINTcan drain into the UART without dropping.
- Reset both boards, set IPs, start PTP, wait for FINE (~2–3 s).
- Send
ptp_offset_reseton FOL. - Wait
--capture-sseconds with no UART traffic. - Send
ptp_offset_dumpon FOL, parse the output. - Compute statistics per sync_status (UNINIT / MATCHFREQ / HARDSYNC / COARSE / FINE): mean, stdev, min, max, p50, p95, |abs_mean|.
- Optional CSV export for external plotting / Allan-deviation analysis.
On the current firmware (commit c0c0be4 with EIC ISR anchor + async
timeout) over a 60 s capture window (~469 FINE-state samples):
Status count mean stdev min max p50 p95 |abs_mean|
FINE 469 -45 24 -201 +69 -47 +2 46 (ns)
Key numbers:
|offset|mean : 46 ns- stdev : 24 ns
- p95 : +2 ns (95 % of samples within ±50 ns)
- worst case : 201 ns over 60 s
This resolves ~550× better than the CLI-based tests (§5.7, §5.8, both showed ~100 µs stdev), and proves the underlying PTP synchronisation is truly sub-100-ns — professional-grade over a 10BASE-T1S multi-drop bus.
python ptp_offset_capture.py --gm-port COM8 --fol-port COM10
python ptp_offset_capture.py --gm-port COM8 --fol-port COM10 --capture-s 120
python ptp_offset_capture.py --gm-port COM8 --fol-port COM10 --csv offsets.csvThe 1024-sample ring buffer wraps after ~128 s at the default 125 ms Sync
period. Use --capture-s ≤ 120 for a complete, non-wrapped trace; longer
runs are supported too but the overwrites counter in the dump header
indicates how many oldest samples were discarded.
Imports helpers from ptp_drift_compensate_test.py for serial handling,
so all three scripts (ptp_drift_compensate_test.py,
ptp_sync_before_after_mux_test.py, ptp_offset_capture.py) must live
in the same directory.
Complements §5.9 from the opposite direction. While §5.9 measures the best-case sync accuracy exploitable when hardware timestamping is available (~50 ns), §5.10 measures the accuracy a pure-software sync protocol would achieve on exactly the same hardware — i.e. the noise floor that remains when every timestamp is taken in application code, after the Harmony TCP/IP stack and SPI transfers.
A minimal NTP-style request/response protocol is implemented in
src/sw_ntp.c (UDP, port 12345, 32-byte packet, four timestamps T1–T4 all from
PTP_CLOCK_GetTime_ns()). The follower does not discipline the clock — it
only measures. Offsets are accumulated into a 1024-entry ring buffer
(src/sw_ntp_offset_trace.c) and dumped in one batch after the capture window,
so the UART plays no role in the measurement path.
See documentation/ptp/ntp_reference.md for the module-level documentation (protocol, CLI, firmware flow).
- Reset both boards, configure IPs.
- Seed
PTP_CLOCKon both sides in parallel threads viaclk_set 0(necessary becausePTP_CLOCK_GetTime_ns()returns 0 on an un-anchored clock). - Start SW-NTP master on the GM side, follower on the FOL side.
- Phase A — HW-PTP OFF: capture 60 s. The PTP clocks free-run on their crystals; drift dominates.
- Enable HW-PTP master/follower, wait for FINE.
- Phase B — HW-PTP ON (FINE): capture 60 s. HW-PTP holds the PTP clocks in sync; only SW-NTP measurement noise remains.
- Print classical (mean/stdev) and robust (median/MAD/IQR) statistics for each phase, and a side-by-side comparison with linear-regression slope (interpretable as crystal drift in ppm).
| Metric | Phase A (HW-PTP off) | Phase B (HW-PTP on) | Ratio |
|---|---|---|---|
| Valid samples / Timeouts | 54 / 7 | 59 / 2 | |
| Slope (crystal drift) | +165 ppm | +0.09 ppm | 1800× |
| Median offset | +4 502 µs | −150 µs | 30× |
| Robust stdev (1.4826·MAD) | 2 643 µs | 24 µs | 110× |
| Residual robust stdev | 850 µs | 23 µs | 37× |
| Classical stdev | 2 678 µs | 1 106 µs† |
† Classical Phase B stdev is inflated by 2–3 outliers (min −6.4 ms, max +5.5 ms out of 59 samples). A heavy-tail warning fires automatically when the classical stdev exceeds 5× the robust stdev — trust the robust number in that case.
-
Crystal drift on this specific pair: +165 ppm measured from Phase A slope. (Individual SAME54 crystal tolerance ±20–50 ppm each → up to ±100 ppm combined; these boards sit at the upper end.)
-
HW-PTP reduces drift by ~1800×: 165 ppm → 0.09 ppm.
-
Surprise: HW-PTP also reduces short-term jitter by ~37× (residual robust stdev 850 µs → 23 µs). A naive prediction would be that the jitter floor is set by SPI + FreeRTOS + stack latencies, which HW-PTP cannot affect. That prediction is wrong. The PI servo applies many tiny corrections per second, damping the high-frequency crystal noise that would otherwise leak into every SW-NTP reading. The regulated
PTP_CLOCKis not just more accurate; it is inherently steadier. -
Systematic bias ≈ −150 µs in Phase B. Not clock skew — the HW-PTP offsets themselves are in the tens of ns. This is TX/RX stack-path asymmetry: the forward (follower-to-master) and backward (master-to-follower) application- layer latencies differ by ~300 µs, and the NTP formula assumes they are equal. No filter removes this.
-
SW-NTP jitter floor for practical use: ~25 µs RMS when riding on top of HW-PTP. Without HW-PTP, SW-NTP alone cannot resolve sub-millisecond sync on this platform.
python sw_ntp_vs_ptp_test.py --gm-port COM8 --fol-port COM10
python sw_ntp_vs_ptp_test.py --capture-s 120 --poll-ms 500
python sw_ntp_vs_ptp_test.py --csv-a phase_a.csv --csv-b phase_b.csv
python sw_ntp_vs_ptp_test.py --skip-phase-a # only HW-PTP-on runImports Logger, open_port, send_command, wait_for_pattern and regex
constants from ptp_drift_compensate_test.py, and zero_both_clocks from
hw_timer_sync_test.py, so all three scripts must live in the same directory.
5.11 tfuture Coordinated Firing + Crystal Deviation Analysis — tfuture_sync_test.py / tfuture_quick_check.py
Caps the time-sync chain on the application side: §5.9 measures how well the PHY hardware aligns clocks (50 ns), §5.10 measures how much of that is observable from plain application-layer code (25 µs SW-NTP floor), and §5.11 measures how precisely the application can act on the synchronised clock by scheduling a coordinated firing event at an absolute PTP_CLOCK time.
The full documentation/features/tfuture.md covers the module, the four diagnostic scripts,
and the bias-investigation history. Summary here:
Firmware (src/tfuture.c) exposes CLI tfuture_at <target_ns>. When two
HW-PTP-synchronised boards arm the same target, each fires when its own
PTP_CLOCK_GetTime_ns() reaches that value. Hybrid precision: main-loop
polling when > 1 ms away, tight busy-wait-spin within the last 1 ms, so
firing precision ≈ TC0 tick (~17 ns) without needing a TC compare ISR.
A 256-entry ring buffer records (target, actual) pairs for statistical
analysis; dump over UART happens outside the measurement path.
tfuture_sync_test.py— baseline regression, 20 rounds × 2 s lead, full setup (~2 min).tfuture_quick_check.py— fast iteration tool, 10 rounds × 2 s lead, ~40 s with reset or ~25 s with--no-reset. Emits a one-line PASS/FAIL verdict and the crystal-deviation summary (see below).tfuture_diagnose_test.py— 4-phase lead-ms scan + drift-toggle, ~4 min. Used when you want to confirm proportional-vs-fixed bias behaviour.tfuture_anchor_delay_test.py/tfuture_drift_forced_test.py/tfuture_drift_forced_fol_test.py— historical sweep tools used during the bias investigation; kept as regression probes.
After both bias fixes (see documentation/features/tfuture.md §8 and the git log of
fix(ptp_clock) + fix(ptp_fol) commits):
| Metric | Median | Robust stdev |
|---|---|---|
| GM self_jitter @ lead=2 s | +9 µs | 60 µs |
| FOL self_jitter @ lead=2 s | +2 µs | 50 µs |
| Inter-board @ lead=2 s | +28 µs | 100 µs |
"Inter-board" is the physical coincidence of the two boards' firings, derived from each board's self-reported PTP_CLOCK value at the firing moment. Sub-100 µs inter-board coincidence is the floor set by the HW-PTP sync accuracy; the tfuture module itself contributes < 50 µs self-jitter per board.
tfuture_quick_check.py also derives the ppm deviation of all four
crystals in the system (2× SAME54 for TC0, 2× LAN8651 for TSU) relative
to GM_LAN8651 as reference. Works because:
drift_ppb_GMdirectly gives GM_SAME54 deviation.drift_ppb_FOLdirectly gives FOL_SAME54 deviation (since the PI servo regulates FOL's TSU rate to match GM's TSU).- Reading the live
MAC_TI+MAC_TISUBNregisters on both boards vialan_readgives the ratio of CLOCK_INCREMENT values, from which FOL_LAN8651's deviation is computed. - GM_LAN8651 is the reference by choice.
Example output on this board pair:
Crystal deviations (reference: GM_LAN8651 = 0 ppm)
GM LAN8651 : 0 ppm (reference)
GM SAME54 : -1028 ppm (from GM drift_ppb)
FOL SAME54 : -1012 ppm (PI makes FOL_TSU = GM_TSU)
FOL LAN8651 : -5 ppm (from live CLOCK_INCREMENT ratio)
The ~1100 ppm mismatch between LAN8651 and SAME54 crystals on each board
is what originally caused the catastrophic ~1.3 ms tfuture bias — the old
DRIFT_SANITY_PPB_ABS = 200 000 (±200 ppm) sanity clamp silently rejected
every sample and the drift filter never converged. See documentation/features/tfuture.md
§8 for the full story.
python tfuture_sync_test.py --gm-port COM8 --fol-port COM10
python tfuture_quick_check.py --gm-port COM8 --fol-port COM10
python tfuture_quick_check.py --no-reset # repeat in ~25 s
python tfuture_diagnose_test.py # full 4-phase when neededAll tfuture scripts import helpers from ptp_drift_compensate_test.py, so
they must live in the same directory.
§5.12 extends §5.11's single-shot firing into a periodic GPIO action on
PD10 of both boards, scheduled entirely from the PTP wallclock. Given
two PTP-locked boards armed with the same period_us and phase_anchor_ns,
both boards drive their PD10 pins at identical PTP moments — scope-verifiable
synchronous signals across the link.
Default period: 1000 µs (full rectangle period) → 1 kHz. The argument
is the full rectangle period; the internal callback fires twice per period
(every period_us / 2) to keep edge jitter bounded.
Firmware (src/cyclic_fire.c, src/cyclic_fire_cli.c):
- Uses
tfuture_set_fire_callback()to register a post-fire hook. - Each hook call acts on PD10 according to the selected pattern (see below),
then calls
tfuture_arm_at_ns(target + half_period)to schedule the next callback at an absolute PTP-wallclock time. - Lowers
tfuture_set_spin_threshold_us()to 100 µs on start (and restores on stop) so PTP / TCP-IP still get CPU between fires at sub-ms periods. - Counts cycles + missed slots for diagnostics.
Two output patterns:
| Pattern | Behaviour | CLI |
|---|---|---|
SQUARE (default) |
One toggle per callback → 50/50 square wave | cyclic_start |
MARKER |
10-callback cycle: callback 0 → HIGH, callback 2 → LOW, callbacks 1 + 3..9 → no-op. Result: one rising edge every 5 × period_us, signal HIGH for 1 period, LOW for 4 periods |
cyclic_start_marker |
MARKER is meant for the "who fires first?" diagnostic — an isolated rising edge is unambiguous to read on a scope, whereas in SQUARE mode at sub-100 µs cross-board offsets the edges look flat-topped overlapped.
Usage (run on both boards after PTP FINE):
# On GM: pick an anchor 2 s in the future, shared with FOL
> clk_get
clk_get: 42195000000 ns drift=+900000ppb
# On GM and FOL, same command — square-wave pattern for measurement:
> cyclic_start 1000 44195000000
cyclic_start OK period=1000 us anchor=44195000000 ns
# or for visual demo:
> cyclic_start_marker 1000 44195000000
cyclic_start_marker OK period=1000 us anchor=44195000000 ns (1-high + 4-low pattern)
# or for the "before sync" half of the demo (no PTP required):
> cyclic_start_free 1000
cyclic_start_free OK period=1000 us (free-run, no PTP sync)
# Wait some time, then:
> cyclic_status
cyclic running : yes
period : 1000 us
cycles : 14231
misses : 0
# Stop:
> cyclic_stop
cyclic stopped
Verification path:
- Firmware-only regression —
smoke_test.pyPhase 3 starts cyclic on both boards for 2 s and checks that cycles accumulate (no callback-dead bug) and misses stay low (no main-loop starvation). - Physical verification via Saleae automation —
cyclic_fire_hw_test.pydrives the full test: reset → PTP FINE →cyclic_starton both boards with a shared anchor → 3 s Saleae Logic 8 capture (Ch0=GM, Ch1=FOL) → edge extraction → cross-board delta statistics with verification timestamps for Logic-2-cursor cross-check. Run with--markerfor the isolated-pulse visual form. CSV of per-edge deltas written for offline plotting. - Drift-filter characterisation —
drift_filter_analysis.pypollsclk_geton both boards at ~1-2 Hz for 60 s and computes per-boarddrift_ppbstddev, autocorrelation, trend and cross-board rate residual (ppm). Use after firmware changes toDRIFT_IIR_Nor to the anchor-tick capture path.
Freeze-state measurement (2026-04-24, canonical test
pd10_sync_before_after_test.py, 10 s capture each phase @ 50 MS/s,
demo decimator path, default drift_iir_n = 128, ISR-captured GM anchor):
| Metric | UNSYNCED | SYNCED (freeze state) |
|---|---|---|
| Cross-board drift slope | +28.72 ppm | −0.07 ppm |
| Cross-board drift MAD | 70.76 µs | 13.62 µs |
| Start → end | +175 → +545 µs | −51.7 → −3.1 µs |
| Per-board interval MAD @ 500 µs | 0.26-0.28 µs | 0.22-0.28 µs |
| PTP FINE lock time | — | 2.7 s |
Synchronisation reduces cross-board drift MAD by 5.2× and collapses
the secular slope to effectively zero. Both gates from the PASS criteria
(|slope| < 5 ppm, MAD < 50 µs) pass with wide margin.
These numbers improve on the earlier documented state (22.3 µs MAD, −1.2 ppm slope — see documentation/testing/pd10_sync_before_after_tests.md) by ~1.6× on MAD and ~17× on residual slope, with no firmware changes to the PTP servo — improvement comes from the adaptive-IIR drift filter and ISR-anchor path already landed on master.
A simple, PTP-independent rectangle generator on the same PD10 pin
used by cyclic_fire. No tfuture callback, no spin-wait, no servo —
just SYS_TIME_Counter64Get() + SYS_PORT_PinToggle() scheduled from
the main loop. Useful for:
- Scope-probe / wiring verification after a fresh flash — enable
blink, confirm edges appear on your probe, then disable it again. - Quick frequency checks of the MCU-local timer path without needing two boards or PTP to be locked.
- Background reference while measuring other subsystems — e.g.
run
blink 1on GM to produce a 1 Hz tick you can cross-correlate with log events.
Firmware (src/pd10_blink.c/.h, src/pd10_blink_cli.c/.h): the module
stays silent on boot; the pin is output-enabled + driven LOW in
APP_Initialize, no further activity until the CLI enables it.
CLI:
| Command | Effect |
|---|---|
blink |
Start at the default 1000 Hz rectangle on PD10. |
blink <hz> |
Start / retune to <hz> Hz (any positive integer). |
blink 0 or blink stop |
Stop; PD10 goes LOW. |
Example scope-check run (requires Saleae):
# on one board, via its serial console:
> blink
blink: running on PD10 at 1000 Hz
# then on the host:
python saleae_freq_check.py --duration 5 --nominal-hz 1000
Because blink uses only the MCU's SYS_TIME and is not disciplined
by PTP, its frequency will sit ~1000 ppm off nominal due to the SAME54
quartz drift (the crystal-deviation analysis in §5.11 quantifies this).
That offset is exactly what saleae_freq_check.py will report.
Relation to cyclic_fire (§5.12): both generate rectangles on
PD10, but only cyclic_fire is PTP-synchronised across boards. Do
not run both simultaneously — the last one started wins the pin.
- MCU: ATSAME54P20A (Cortex-M4F, 120 MHz)
- Ethernet: LAN865x 10BASE-T1S via SPI (TC6 protocol)
- Framework: MPLAB Harmony 3 (bare-metal, no FreeRTOS)
- Build: CMake 4.1 + Ninja
- HEX output:
tcpip_iperf_lan865x.X/out/tcpip_iperf_lan865x/default.hex - Programmer: MPLAB MDB (
flash.py)
| Board | Role (default) | Serial |
|---|---|---|
| Board 1 | Grandmaster (COM8) | ATML3264031800001049 |
| Board 2 | Follower (COM10) | ATML3264031800001290 |
python setup_compiler.py # select XC32 version (patches toolchain.cmake)
python setup_flasher.py # assign Board 1 / Board 2 to connected debuggers
build.bat # compile (summary printed automatically)
python flash.py # flash both boardsOne-time setup tool — scans C:\Program Files\Microchip\xc32\ for installed
XC32 versions, lets the user pick one, patches all version-string occurrences in
cmake/.generated/toolchain.cmake (both forward-slash and double-backslash
forms, covering all 17 entries: CMAKE_C_COMPILER, CMAKE_AR, MP_BIN2HEX,
etc.), and saves the choice to setup_compiler.config (JSON, git-ignored).
Installed XC32 versions (2 found):
[1] v4.60 C:\Program Files\Microchip\xc32\v4.60\bin\xc32-gcc.exe
[2] v5.10 C:\Program Files\Microchip\xc32\v5.10\bin\xc32-gcc.exe <-- current
[0] Abort / keep current selection
Select version number: 1
...
Patched toolchain.cmake: v5.10 -> v4.60
Done. build.bat will use XC32 v4.60.
Single-command CMake + Ninja build.
| Parameter | Behaviour |
|---|---|
(none) / incremental |
Only recompiles changed files |
clean |
Deletes temporary build directory |
rebuild |
Clean + full build |
help |
Prints options |
Reads setup_compiler.config at startup; aborts with a clear error if the file
is missing or the configured xc32-gcc.exe does not exist.
CMake intermediate files (.o, .d, build.ninja) are placed outside the
repository at C:\work\ptp\AN1847\harmony\temp\tcpip_iperf_lan865x\default\
to avoid Windows MAX_PATH (260 character) issues. The path is derived
automatically at runtime relative to build.bat — no hardcoded absolute path,
works on any machine (provided the project is checked out inside a harmony\
parent directory).
Called automatically by build.bat after every successful build. Parses the
linker output and ELF symbol table to produce a concise human-readable summary.
| Source | Information extracted |
|---|---|
memoryfile.xml |
Flash used/free/total, RAM used/free/total |
mem.map |
_min_heap_size (heap reserved by linker script) |
default.elf via xc32-nm |
Active interrupt handler names (weak Dummy_Handler symbols silently skipped) |
default.elf binary scan |
Build timestamp (__DATE__ / __TIME__ embedded by app.c) |
Example output:
==============================================================
BUILD SUMMARY
==============================================================
Build : Apr 8 2026 17:08:51
Flash (program memory)
Used : 131,877 bytes ( 128.8 KiB) 12.6%
Free : 916,699 bytes ( 895.2 KiB)
[####--------------------------]
RAM (data memory)
Used : 15,937 bytes ( 15.6 KiB) 6.1%
Free : 246,207 bytes ( 240.4 KiB)
[##----------------------------]
Linker-Reserved Regions
Heap : 44,960 bytes ( 43.9 KiB) (_min_heap_size)
Interrupt Handlers
Core IRQs ( 7): BusFault, DebugMonitor, HardFault,
MemoryManagement, NonMaskableInt, Reset, UsageFault
Peripheral IRQs ( 5): DMAC_0, DMAC_1, SERCOM0_SPI, SERCOM1_USART, TC0_Timer
Image HEX : ...image\tcpip_iperf_lan865x_20260408_170851.hex
Summary : ...image\build_summary_20260408_170851.txt
Versioned artefacts are written to out/tcpip_iperf_lan865x/image/ and tracked
by git (!**/image/*.hex negation rule in .gitignore) so released binaries
are available after git clone without a rebuild.
Detects connected EDBG debuggers (USB VID 0x03EB, serial number prefix ATML,
or manufacturer string containing microchip / atmel), lets the user assign
Board 1 (Grandmaster) and Board 2 (Follower), saves to setup_flasher.config
(JSON, git-ignored).
Flash default.hex to one or both boards via MPLAB MDB. Board serial numbers
and COM ports are read from setup_flasher.config.
python flash.py [--board1-only | --board2-only] [--hex <path>] [--swd-khz <n>]Workaround for a CMake + xc32-gcc (MINGW) incompatibility: MINGW strips
backslashes from command-line arguments, so the linker receives
-o out\default.elf and creates a file literally named outdefault.elf — the
bin2hex step then fails with No such file or directory.
user.cmake redirects the linker output to an absolute forward-slash path
(unaffected by the MINGW bug) and adds a POST_BUILD copy step to the canonical
location out/tcpip_iperf_lan865x/default.elf. Loaded via the standard
include(user.cmake OPTIONAL) hook in CMakeLists.txt — no manual action
required.
The project can also be opened and built directly in MPLAB X IDE as an
alternative to build.bat. The nbproject/ directory contains all necessary
project metadata; MPLAB X generates the required Makefile-impl.mk and
Makefile-variables.mk automatically when the project is opened.
- Open MPLAB X IDE
- File → Open Project → select
tcpip_iperf_lan865x.X/ - MPLAB X detects the project type (
com.microchip.mplab.nbide.embedded.makeproject) and generates the missing Makefile fragments automatically - Select the
defaultconfiguration - Production → Build Main Project (or press F11)
- The XC32 toolchain version configured in MPLAB X must match the version used
by
setup_compiler.py/build.batto ensure identical compiler flags - All PTP source files (
ptp_clock.c,ptp_gm_task.c,ptp_fol_task.c,filters.c) are registered innbproject/configurations.xmland will appear in the MPLAB X project tree automatically - MPLAB X uses its own intermediate build directory (inside
build/) — this is separate from the CMake build tree underC:\...\temp\ - For flashing,
flash.py(MPLAB MDB) remains the recommended method; MPLAB X can also program via Run → Run Main Project if a debugger is attached
The project can be opened and built in Visual Studio Code using the pre-configured CMake integration. VS Code acts as the editor and build frontend; the actual compiler is still XC32 via CMake + Ninja.
- CMake Tools extension
- C/C++ extension (for IntelliSense)
setup_compiler.pyrun at least once (createssetup_compiler.config)
- Open VS Code in the project working directory:
(from
code .
tcpip_iperf_lan865x.X\) - VS Code detects the
cmake/folder and the CMake Tools extension activates automatically - Select the default configure preset when prompted
- Build: press
Ctrl+Shift+Bor use the CMake Tools status bar button (▶ Build)
- Build output (
.o,.elf,.hex) goes to the same location asbuild.bat:C:\...\temp\tcpip_iperf_lan865x\default\ build_summary.pyis not called automatically from VS Code — runpython build_summary.pymanually after a build if needed- IntelliSense uses
compile_commands.jsongenerated by CMake in the temp build directory; the C/C++ extension picks it up automatically once CMake has configured at least once - For flashing, run
python flash.pyfrom the integrated terminal
The repository includes a pre-configured launch.json with two debug
configurations:
| Config | Description |
|---|---|
| Launch tcpip_iperf_lan865x: default | Builds, programs, and starts the debugger in one step |
| Attach (after flash.py) | Attaches to an already-flashed device (workaround for tool pack bugs) |
One-time prerequisite: run setup_debug.py once per machine after clone:
python setup_debug.pyThis patches %USERPROFILE%\.mchp_packs\Microchip\SAME54_DFP\<ver>\scripts\dap_cortex-m4.py
to add a missing global variable is_debug_build that tool pack version 1.6.762
omits. Without this patch, every debug session fails at the Erasing... step with:
NameError: global name 'is_debug_build' is not defined
Failed to start session: Debugger::program : Failed to program the target device
The fix is idempotent — running setup_debug.py multiple times is safe.
This chapter covers two related but distinct concepts:
- Coding with AI: GitHub Copilot / Claude write the orchestrator code — this happens during development and is already standard practice in this project.
- RL in the loop: An algorithm (ranging from a simple for-loop up to an LLM) selects the next firmware parameters at runtime, builds the firmware, flashes the board, and interprets the measurement as a reward signal.
The core argument: the hardware-in-the-loop environment is already fully present in this project. Only the orchestrator that drives it is missing.
Reinforcement Learning (RL) is a framework in which an Agent repeatedly takes Actions, observes the resulting State of an Environment, and receives a scalar Reward signal. The agent's goal is to maximise cumulative reward over time.
Applied to embedded firmware development the loop looks like this:
┌─────────────────────────────────────────────────────────────┐
│ AGENT (see §7.4 for three levels) │
│ │
│ Level 1 — for-loop / grid search (no AI model) │
│ Level 2 — Bayesian optimiser (statistical model) │
│ Level 3 — LLM as policy (real AI API call) │
└───────────────────┬──────────────────────▲──────────────────┘
│ Action │ Reward + next State
│ (new #define values) │ (stdev, fine_s, pass)
▼ │
┌─────────────────────────────────────────────────────────┐
│ ENVIRONMENT │
│ │
│ 1. render_params() — writes #define values to params.h│
│ 2. build.bat — compiles → hex file │
│ 3. flash.py — programs the board │
│ 4. ptp_reproducibility_test.py — measures, emits JSON │
│ • stdev_ns (lower = better) │
│ • slope_ppm (closer to 0 = better) │
│ • fine_s (lower = better) │
│ • pass (boolean) │
└─────────────────────────────────────────────────────────┘
Key point: The Environment is identical across all three levels. Only the Agent changes.
The environment already exists in this project:
| Environment step | Tool already present |
|---|---|
| Edit firmware | ptp_fol_task.h, filters.h (plain #define values) |
| Build | build.bat (one-command, reproducible) |
| Flash | flash.py (one-command, automatic port detection) |
| Measure | ptp_reproducibility_test.py (structured JSON output) |
Only the orchestrator layer is missing — a script that strings these four steps together and drives a search algorithm or AI model over the parameter space.
Regardless of the chosen agent level, the orchestrator must:
-
Parameterise the firmware — substitute numeric
#definevalues before each build.
Safest approach: aparams_template.hwith{{PLACEHOLDER}}tokens; the orchestrator renders it into the real header before callingbuild.bat. -
Build and flash without operator interaction — both tools already support this; the orchestrator calls them as subprocesses and checks the exit code.
-
Parse the reward signal —
ptp_reproducibility_test.pyhas a--jsonoutput mode; the orchestrator reads the resulting file. -
Implement a search or learning policy — Level 1: for-loop. Level 2: Bayesian optimiser. Level 3: LLM (see §7.4).
-
Gate dangerous actions — optionally require a
y/nconfirmation before flashing if the proposed parameter value is outside a safe operating envelope (e.g. a TI value that would exceed hardware limits).
Minimal orchestrator skeleton (shared by all three levels):
import subprocess, json, pathlib
PARAMS_HEADER = pathlib.Path("src/params.h") # re-created before every build
def render_params(params: dict) -> None:
"""Write a params.h containing the desired #define values."""
lines = [f"#define {k} {v}" for k, v in params.items()]
PARAMS_HEADER.write_text("\n".join(lines) + "\n")
def build() -> bool:
return subprocess.run(["build.bat"], check=False).returncode == 0
def flash() -> bool:
return subprocess.run(["python", "flash.py"], check=False).returncode == 0
def measure(port_gm: str, port_fol: str) -> dict:
result = subprocess.run(
["python", "ptp_reproducibility_test.py",
"--gm", port_gm, "--fol", port_fol, "--json", "result.json"],
check=False,
)
if result.returncode != 0:
return {"pass": False, "stdev_ns": 1e9, "slope_ppm": 1e6, "fine_s": 1e6}
return json.loads(pathlib.Path("result.json").read_text())
def reward(metrics: dict) -> float:
"""Reward function: lower stdev and faster FINE lock = higher reward."""
if not metrics["pass"]:
return -1000.0
return -metrics["stdev_ns"] - 0.1 * metrics["fine_s"]
def run_episode(params: dict, port_gm: str, port_fol: str) -> float:
"""One full cycle: parameters → build → flash → measure → reward."""
render_params(params)
if not build() or not flash():
return -1000.0
return reward(measure(port_gm, port_fol))The servo has several numeric constants that are good candidates for automated optimisation. All live in plain C headers — no MCC regeneration required.
| Parameter | Location | Current value | Effect |
|---|---|---|---|
PTP_SYNC_INTERVAL |
ptp_fol_task.h:80 |
500 ms | Sync message period; lower = faster reaction, higher = smoother frequency estimate |
FIR_FILER_SIZE |
filters.h:40 |
16 taps | Rate-ratio FIR smoothing; larger = less noise, more lag |
FIR_FILER_SIZE_FINE |
filters.h:41 |
3 taps | Offset FIR in COARSE/FINE state; trades jitter vs. responsiveness |
HARDSYNC_COARSE_THRESHOLD |
ptp_fol_task.h:86 |
300 ns | Offset boundary HARDSYNC→COARSE; too small = coarse never reached; too large = noise triggers coarse |
HARDSYNC_FINE_THRESHOLD |
ptp_fol_task.h:87 |
150 ns | Offset boundary COARSE→FINE; governs when TISUBN fine-tuning activates |
MATCHFREQ_RESET_THRESHOLD |
ptp_fol_task.h:83 |
100 000 000 ns | Safety guard: offset above this resets to MATCHFREQ |
Parameters interact non-linearly: a smaller HARDSYNC_FINE_THRESHOLD makes FINE reachable faster but requires a correspondingly small FIR_FILER_SIZE_FINE to avoid oscillation. This coupling is exactly why manual hand-tuning is tedious and RL is attractive.
PTP_SYNC_INTERVAL is a 1-D optimisation problem — a good first test for the infrastructure. The three levels below show how the agent becomes progressively smarter while the Environment remains unchanged.
Reward function (identical for all three levels):
No AI call. The agent is a for-loop;
run_episodeis the only thing invoked.
# grid_search_interval.py
import argparse, json, pathlib
from orchestrator import run_episode
CANDIDATES = [62, 125, 250, 500, 1000] # milliseconds
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--gm", required=True)
ap.add_argument("--fol", required=True)
args = ap.parse_args()
results = []
for interval in CANDIDATES: # ← this IS the entire "agent"
params = {"PTP_SYNC_INTERVAL": f"{interval}u"}
r = run_episode(params, args.gm, args.fol) # build → flash → measure
print(f"interval={interval:5d} ms reward={r:10.1f}")
results.append({"interval_ms": interval, "reward": r})
best = max(results, key=lambda x: x["reward"])
print(f"\nBest: {best['interval_ms']} ms (reward {best['reward']:.1f})")
pathlib.Path("grid_results.json").write_text(json.dumps(results, indent=2))
if __name__ == "__main__":
main()Runtime: 5 × (≈30 s build/flash + ≈30 s test) ≈ 5 minutes — a full characterisation in the time it takes to make a coffee.
Expected output (values illustrative):
interval= 62 ms reward= -87.3
interval= 125 ms reward= -45.1
interval= 250 ms reward= -38.9
interval= 500 ms reward= -42.2
interval= 1000 ms reward= -61.7
Best: 250 ms (reward -38.9)
No LLM call. A Gaussian process model (
skopt) learns a surrogate function over the parameter space and selects the next candidate so as to maximise information gain. This significantly reduces the number of required build-flash cycles — important when the search space is larger (multiple parameters simultaneously).
# bayes_search_interval.py
import argparse
from skopt import gp_minimize
from skopt.space import Integer
from orchestrator import run_episode
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--gm", required=True)
ap.add_argument("--fol", required=True)
ap.add_argument("--calls", type=int, default=12) # total number of builds
args = ap.parse_args()
def objective(params):
interval = params[0]
r = run_episode({"PTP_SYNC_INTERVAL": f"{interval}u"}, args.gm, args.fol)
print(f"interval={interval:5d} ms reward={r:10.1f}")
return -r # skopt minimises → negate
result = gp_minimize(
objective,
dimensions=[Integer(62, 1000, name="interval_ms")],
n_calls=args.calls, # ← Bayesian model picks every next point
random_state=42,
)
print(f"\nBest interval: {result.x[0]} ms (reward {-result.fun:.1f})")
if __name__ == "__main__":
main()With --calls 12, Bayesian optimisation typically matches the result of a full grid search over 20–30 points. The benefit grows significantly when tuning multiple parameters simultaneously (§7.3).
This is where AI is called. The LLM sees the accumulated measurement history and proposes the next candidate — as a chat prompt. It acts as an "intelligent" agent that can also provide reasoning.
# llm_search_interval.py
import argparse, json
import openai
from orchestrator import run_episode
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--gm", required=True)
ap.add_argument("--fol", required=True)
ap.add_argument("--rounds", type=int, default=8)
args = ap.parse_args()
client = openai.OpenAI() # API key from environment variable OPENAI_API_KEY
history = []
for round_nr in range(args.rounds):
# ── Step 1: ask the LLM which value to test next ─────────────────────
system_prompt = (
"You are an optimisation assistant for a PTP timestamp servo. "
"Your task: choose the next value for PTP_SYNC_INTERVAL (ms) "
"to maximise the reward r = -stdev_ns - 0.1*fine_s. "
"Allowed values: 62..2000 ms (integer). "
"Reply with a single integer only, no explanation."
)
user_msg = (
f"Results so far: {json.dumps(history, indent=2)}\n\n"
f"Round {round_nr + 1}/{args.rounds}: which value should I test next?"
)
response = client.chat.completions.create( # ← AI API call
model="gpt-4o",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_msg},
],
max_tokens=10,
temperature=0.2,
)
interval = int(response.choices[0].message.content.strip())
print(f"[LLM suggests] interval={interval} ms")
# ── Step 2: build, flash, measure ────────────────────────────────────
r = run_episode({"PTP_SYNC_INTERVAL": f"{interval}u"}, args.gm, args.fol)
history.append({"round": round_nr + 1, "interval_ms": interval, "reward": r})
print(f" → reward={r:.1f}")
best = max(history, key=lambda x: x["reward"])
print(f"\nBest result: {best['interval_ms']} ms (reward {best['reward']:.1f})")
if __name__ == "__main__":
main()Example console output:
[LLM suggests] interval=250 ms
→ reward=-41.3
[LLM suggests] interval=200 ms
→ reward=-38.1
[LLM suggests] interval=175 ms
→ reward=-36.8
[LLM suggests] interval=150 ms
→ reward=-39.7
...
Best result: 175 ms (reward -36.8)
After each measurement the LLM sees the complete history JSON and can explain its next suggestion — e.g. "175 ms was better than 200 ms, so I'll try 160 ms next". This explainability distinguishes Level 3 from Level 2.
| Level | Agent | AI model | Builds for a good result | Setup required |
|---|---|---|---|---|
| 1 — Grid Search | for-loop | none | N (all candidates) | no external library |
| 2 — Bayesian Opt. | skopt GP model |
statistical | ~12–15 | pip install scikit-optimize |
| 3 — LLM Policy | GPT-4o / Copilot | LLM (API call) | ~8–12 | OpenAI API key |
Note on "Coding with AI": The orchestrator scripts, the
run_episodeskeleton, and this entire chapter were written with the help of GitHub Copilot (Claude Sonnet). The workflow — Copilot proposes code → firmware is built and measured → results feed back into the next prompt — is itself the closed loop described in §7.1, applied at the level of the development process rather than inside the firmware.
Several Python scripts in this repository require third-party packages (e.g. pyserial).
Two helper files automate the detection and installation of these packages.
| File | Purpose |
|---|---|
analyze_dependencies.py |
Scans all .py files, detects third-party imports, writes requirements.txt |
install_dependencies.bat |
Windows batch script that installs every package listed in requirements.txt |
requirements.txt |
Auto-generated list of pip packages required by this repository |
Step 1 — Analyze (run once, or after adding new Python scripts)
python analyze_dependencies.py
The script walks the entire repository, parses every .py file with Python's ast
module, filters out standard-library modules and local files, and writes a fresh
requirements.txt.
Step 2 — Install (Windows)
Double-click install_dependencies.bat or run it from a command prompt:
install_dependencies.bat
The batch script:
- Checks that Python is installed and available on
PATH - Checks that
pipis available - Upgrades pip to the latest version
- Runs
pip install -r requirements.txt - Shows clear status messages and error hints if anything goes wrong
- Python 3.8 or newer — https://www.python.org/downloads/ (check "Add Python to PATH" during installation)
- An internet connection for the initial package download
This chapter provides a detailed, file-level analysis of the PTP (IEEE 1588) implementation: Grandmaster (GM) and Follower (FOL), the complete synchronisation message flow, and annotated pseudo-code.
| File | Role |
|---|---|
src/ptp_gm_task.c |
GM state machine — sends Sync/FollowUp, reads TX timestamps |
src/ptp_gm_task.h |
GM public API, register macros, timing constants |
src/ptp_fol_task.c |
FOL message processing, clock servo algorithm |
src/ptp_fol_task.h |
FOL public API, all PTP wire-format structs |
src/ptp_clock.c/.h |
Software PTP wallclock (ns resolution, TC0 interpolation) |
src/ptp_ts_ipc.h |
IPC structs for HW RX-timestamps between driver and app |
src/app.c |
Application state machine — service dispatcher for GM and FOL |
src/config/default/driver/lan865x/src/dynamic/drv_lan865x_api.c |
LAN865x driver, TC6_CB_OnRxEthernetPacket() |
PTP_GM_Init() is called when the user issues ptp_mode master
(app.c:167, ptp_gm_task.c:349).
PTP_GM_Init()
→ read board MAC via TCPIP_STACK_NetAddressMac()
→ adopt calibrated TI/TISUBN values from FOL servo (if available)
→ enter state GM_STATE_RMW_CONFIG0_READ
Pre-init RMW sequence (mirrors TC6_ptp_master_init steps 8 and 9):
- OA_CONFIG0 RMW (
0x00000004): set bits 7 and 6 (FTSE) — enables hardware timestamping (ptp_gm_task.c:729–777, mask0xC0). - PADCTRL RMW (
0x000A0088): set bit 8, clear bit 9 (ptp_gm_task.c:779–829).
Normal init sequence — 9 sequential register writes
(ptp_gm_task.c:157–173):
| Register | Address | Value | Purpose |
|---|---|---|---|
GM_TXMCTL |
0x00040040 |
0x0000 |
Reset TX-Match detector |
GM_TXMLOC |
0x00040045 |
30 |
Match position inside frame |
GM_TXMPATH |
0x00040041 |
0x88 |
Pattern: EtherType high byte |
GM_TXMPATL |
0x00040042 |
0xF710 |
Pattern: EtherType low byte + offset |
GM_TXMMSKH/L |
0x00040043/44 |
0x00 |
Pattern masks |
MAC_TISUBN |
0x0001006F |
calibrated sub-increment | Sub-nanosecond clock increment |
MAC_TI |
0x00010077 |
calibrated TI (default: 40) | Nanosecond clock increment |
PPSCTL |
0x000A0239 |
0x7D |
Enable 1PPS output |
PTP_GM_Service() is called every 1 ms from APP_STATE_IDLE (app.c:489).
The full per-cycle sequence:
GM_STATE_WAIT_PERIOD ── every 125 ms ──▶ GM_STATE_SEND_SYNC
build_sync()
WriteRegister(GM_TXMCTL, 0x0002) // TXME=1: arm TX-Match detector
wait write callback
SendRawEthFrame(sync_buf, tsc=1) // tsc=1 → capture TX timestamp
wait TX-done callback
→ GM_STATE_READ_STATUS0
DRV_LAN865X_GetAndClearTsCapture() // check TTSCAA/B/C
ReadRegister(GM_OA_TTSCAH) // t1 seconds
ReadRegister(GM_OA_TTSCAL) // t1 nanoseconds
WriteRegister(GM_OA_STATUS0, status0) // W1C: clear capture flags
→ GM_STATE_SEND_FOLLOWUP
build_followup(t1_sec, t1_nsec + 7650) // apply PTP_GM_STATIC_OFFSET
SendRawEthFrame(followup_buf, tsc=0)
wait TX-done callback
PTP_CLOCK_Update(t1 + GM_ANCHOR_OFFSET_NS, SYS_TIME_Counter64Get())
gm_seq_id++
→ GM_STATE_WAIT_PERIOD
Sync frame layout (build_sync, ptp_gm_task.c:263–289):
- 14-byte Ethernet header — EtherType
0x88F7, destination broadcast or PTP L2 multicast01:80:C2:00:00:0E - 44-byte
syncMsg_t—tsmt=0x10,version=0x02,messageLength=0x002C,sequenceID,flags[0]=0x02, flags[1]=0x08 originTimestampis zero — the precise value is carried by the FollowUp
FollowUp frame layout (build_followup, ptp_gm_task.c:291–319):
preciseOriginTimestamp= TTSCA +PTP_GM_STATIC_OFFSET(7 650 ns TX-path compensation)- Organisation-Specific TLV type
0x0003, OUI00:80:C2, sub-type01(cumulative rate-ratio)
// ptp_gm_task.c:70–97 — runtime state
gmState_t gm_state; // current state-machine state
uint32_t gm_ts_sec/nsec; // TX timestamp from TTSCA registers
uint16_t gm_seq_id; // PTP sequence ID
uint32_t gm_sync_interval_ms; // default 125 ms
// frame buffers
uint8_t gm_sync_buf[60]; // 14 (ETH) + 44 (syncMsg_t) + 2 pad
uint8_t gm_followup_buf[90]; // 14 (ETH) + 76 (followUpMsg_t)PTP_FOL_Init() is called from APP_STATE_IDLE on first entry (app.c:412)
and again from resetSlaveNode() on every follower reset
(ptp_fol_task.c:638):
PTP_FOL_Init()
→ WriteRegister(PPSCTL, 0x02) // stop PPS output
→ WriteRegister(SEVINTEN, PPSDONE_Msk) // enable PPS-done interrupt
→ memset(TS_SYNC, 0) // clear timestamp store
→ initialise FIR/IIR filter buffers
Follower operation starts when PTP_FOL_SetMode(PTP_SLAVE) is called
(e.g. ptp_mode follower CLI command, app.c:170), which in turn calls
resetSlaveNode() (ptp_fol_task.c:663–668).
Hardware data flow:
LAN865x SPI footer carries the RTSA timestamp
→ TC6_CB_OnRxEthernetPacket() [drv_lan865x_api.c:1372]
checks EtherType == 0x88F7
copies frame → g_ptp_raw_rx.data
g_ptp_raw_rx.rxTimestamp = RTSA (sec[63:32] | ns[31:0])
if SYNC (rxTimestamp != NULL):
g_ptp_raw_rx.sysTickAtRx = SYS_TIME_Counter64Get()
g_ptp_raw_rx.pending = true
app.c:505 (APP_STATE_IDLE polling loop)
→ PTP_FOL_OnFrame(data, length, rxTimestamp)
PTP_FOL_OnFrame() [ptp_fol_task.c:688]
→ extracts sec/nsec from rxTimestamp
→ handlePtp(pData, len, sec, nsec)
handlePtp() [ptp_fol_task.c:614]
→ parses messageType = ptpHeader.tsmt & 0x0F
→ MSG_SYNC → processSync() + stores TS_SYNC.receipt = {sec, nsec}
→ MSG_FOLLOW_UP → processFollowUp() (servo core)
processSync() (ptp_fol_task.c:369) validates the sequence ID and stores t2
(the RTSA hardware receive timestamp) in TS_SYNC.receipt.
processFollowUp() (ptp_fol_task.c:392) is the clock-servo core:
// t1 — precise GM send time from FollowUp frame
t1 = preciseOriginTimestamp (byte-swapped) + correctionField / 65536
// t2 — local receive time from RTSA (stored during processSync)
t2 = TS_SYNC.receipt
// update the software clock anchor
PTP_CLOCK_Update(t2, g_ptp_raw_rx.sysTickAtRx)
// frequency error measurement
diffLocal = t2_now - t2_prev // local interval between two SYNCs
diffRemote = t1_now - t1_prev // GM interval between two SYNCs
rateRatio = diffRemote / diffLocal
rateRatioFIR = FIR_filter(rateRatio) // averaged over 16 samples
// offset calculation
offset = t2 - t1 // positive: follower is ahead; negative: follower lagsNote on delay measurement: no Delay_Req / Delay_Resp exchange is implemented. The system uses one-way hardware timestamping only. Propagation delay is compensated by the static constant
PTP_GM_STATIC_OFFSET = 7 650 nsadded to the FollowUp timestamp on the GM side (ptp_gm_task.h:61).
(ptp_fol_task.c:485–572)
| State | Entry condition | Action |
|---|---|---|
UNINIT |
initial state | accumulate 16 rate-ratio samples |
UNINIT → MATCHFREQ |
after 16 samples | compute TI/TISUBN → FOL_ACTION_SET_CLOCK_INC |
MATCHFREQ |
|offset| > 100 000 000 ns |
hardResync = 1 |
MATCHFREQ → HARDSYNC |
|offset| ≤ 100 000 000 ns |
— |
HARDSYNC |
|offset| > 0x3FFF FFFF |
→ UNINIT (full reset) |
HARDSYNC |
|offset| > 16 777 215 ns |
capped direct adjust → FOL_ACTION_ADJUST_OFFSET |
HARDSYNC → COARSE |
|offset| > 300 ns |
FIR-coarse filter → FOL_ACTION_ADJUST_OFFSET |
COARSE/FINE |
|offset| > 150 ns |
FIR-fine filter → FOL_ACTION_ADJUST_OFFSET → FINE |
When hardResync == 1 the follower writes t1 directly into MAC_TSL /
MAC_TN to hard-set the LAN865x hardware clock
(FOL_ACTION_HARD_SYNC, ptp_fol_task.c:421–426).
PTP_FOL_Service() is called every 1 ms (app.c:496). It serialises all
LAN865x SPI writes through the fol_pending_action flag
(ptp_fol_task.c:143–291):
| Action | Registers written | Effect |
|---|---|---|
FOL_ACTION_HARD_SYNC |
MAC_TSL, MAC_TN |
Hard-set LAN865x clock to t1 |
FOL_ACTION_SET_CLOCK_INC |
MAC_TISUBN, MAC_TI |
Apply crystal-drift correction |
FOL_ACTION_ADJUST_OFFSET |
MAC_TA |
Fine-adjust LAN865x clock (signed offset) |
FOL_ACTION_ENABLE_PPS |
PPSCTL |
Enable 1PPS output after first lock |
Timestamp notation
- t1 — precise Sync send time on the GM LAN865x (from TTSCA registers)
- t2 — Sync receive time on the FOL LAN865x (from RTSA in SPI footer)
GRANDMASTER FOLLOWER
(ptp_gm_task.c) (ptp_fol_task.c / drv_lan865x_api.c)
│ │
│ ── every 125 ms ── │
│ │
│ 1. build_sync() │
│ WriteRegister(GM_TXMCTL, 0x0002) │
│ │
│ 2. SendRawEthFrame(sync_buf, tsc=1) │
│ ──── SYNC (EtherType 0x88F7) ────────────────────▶ │
│ LAN865x: captures t1 via TX-Match detector │
│ │ LAN865x: captures t2 via RTSA
│ │ TC6_CB_OnRxEthernetPacket:
│ │ g_ptp_raw_rx.rxTimestamp = t2
│ │ g_ptp_raw_rx.sysTickAtRx = TC0 tick
│ │ g_ptp_raw_rx.pending = true
│ │
│ 3. Read OA_STATUS0 → check TTSCAA │
│ 4. Read TTSCA_H (seconds of t1) │
│ 5. Read TTSCA_L (nanoseconds of t1) │
│ 6. Write OA_STATUS0 (W1C clear) │
│ │
│ 7. build_followup(t1 + 7 650 ns) │
│ SendRawEthFrame(followup_buf, tsc=0) │
│ ──── FOLLOW_UP ───────────────────────────────────▶ │
│ │ app.c: g_ptp_raw_rx.pending:
│ 8. PTP_CLOCK_Update(t1 + anchor_offset, TC0 tick) │ PTP_FOL_OnFrame()
│ (update GM software clock) │ handlePtp → processSync:
│ │ TS_SYNC.receipt = t2
│ │ handlePtp → processFollowUp:
│ │ t1 from preciseOriginTimestamp
│ │ offset = t2 − t1
│ │ rateRatio = diffRemote/diffLocal
│ │ PTP_CLOCK_Update(t2, sysTickAtRx)
│ │ → servo state machine
│ │ → fol_pending_action set
│ │
│ │ PTP_FOL_Service() (1 ms tick):
│ │ WriteRegister(MAC_TSL / MAC_TN)
│ │ or WriteRegister(MAC_TI / TISUBN)
│ │ or WriteRegister(MAC_TA)
│ │
│ ── next 125 ms period ── │
//=======================================================================
// PTP GRANDMASTER — main service loop (called every 1 ms)
//=======================================================================
PROCEDURE PTP_GM_Service():
every 125 ms:
// Step 1 — arm TX-Match detector
WriteRegister(GM_TXMCTL, 0x0002) // TXME = 1
wait write callback
// Step 2 — send SYNC with timestamp-capture request
syncFrame = buildSyncFrame(sequenceID = gm_seq_id)
SendRawEthFrame(syncFrame, tsc = 1)
wait TX callback // frame is now on the wire
// Step 3 — read TX timestamp from LAN865x hardware
wait STATUS0.TTSCAA == 1 // capture slot available
t1_sec = ReadRegister(GM_OA_TTSCAH)
t1_nsec = ReadRegister(GM_OA_TTSCAL)
WriteRegister(GM_OA_STATUS0, status0) // W1C: clear capture flags
// Step 4 — apply static TX-path compensation
t1_nsec += PTP_GM_STATIC_OFFSET // 7 650 ns
if t1_nsec >= 1_000_000_000:
t1_sec += 1
t1_nsec -= 1_000_000_000
// Step 5 — send FOLLOW_UP carrying t1
followUpFrame = buildFollowUpFrame(
sequenceID = gm_seq_id,
preciseOriginTimestamp = { t1_sec, t1_nsec }
)
SendRawEthFrame(followUpFrame, tsc = 0)
wait TX callback
// Step 6 — update software clock anchor
wc_ns = t1_sec * 1_000_000_000 + t1_nsec + GM_ANCHOR_OFFSET_NS
PTP_CLOCK_Update(wc_ns, SYS_TIME_Counter64Get())
gm_seq_id++
//=======================================================================
// FOLLOWER — driver callback (interrupt context)
//=======================================================================
CALLBACK TC6_CB_OnRxEthernetPacket(frame, len, rxTimestamp):
if EtherType == 0x88F7: // PTP frame
g_ptp_raw_rx.data = frame
g_ptp_raw_rx.rxTimestamp = rxTimestamp // t2: RTSA (sec[63:32] | ns[31:0])
if rxTimestamp != NULL: // SYNC only — FollowUp has no TS
g_ptp_raw_rx.sysTickAtRx = SYS_TIME_Counter64Get()
g_ptp_raw_rx.pending = true
//=======================================================================
// FOLLOWER — application task (polling, 1 ms)
//=======================================================================
PROCEDURE APP_STATE_IDLE():
if g_ptp_raw_rx.pending:
g_ptp_raw_rx.pending = false
PTP_FOL_OnFrame(data, len, rxTimestamp)
//=======================================================================
// FOLLOWER — frame entry point
//=======================================================================
PROCEDURE PTP_FOL_OnFrame(pData, len, rxTimestamp):
sec = rxTimestamp[63:32]
nsec = rxTimestamp[31:0]
handlePtp(pData, len, sec, nsec)
PROCEDURE handlePtp(pData, len, sec, nsec):
messageType = pData[14].tsmt & 0x0F // after 14-byte ETH header
if messageType == MSG_SYNC:
processSync(pData)
TS_SYNC.receipt = { sec, nsec } // t2: hardware receive timestamp
elif messageType == MSG_FOLLOW_UP:
processFollowUp(pData)
//=======================================================================
// FOLLOWER — SYNC processing
//=======================================================================
PROCEDURE processSync(syncMsg):
seqId = syncMsg.header.sequenceID (byte-swapped)
if mismatch > 10: resetSlaveNode()
else: syncReceived = 1
//=======================================================================
// FOLLOWER — FOLLOW_UP processing and clock servo
//=======================================================================
PROCEDURE processFollowUp(followUpMsg):
// Extract t1 from wire
t1 = followUpMsg.preciseOriginTimestamp (byte-swapped)
+ correctionField / 65536
// t2 was stored during processSync
t2 = TS_SYNC.receipt
// Update software clock anchor
PTP_CLOCK_Update(t2, g_ptp_raw_rx.sysTickAtRx)
// Frequency-error measurement
diffLocal = t2_now - t2_prev // local interval
diffRemote = t1_now - t1_prev // GM interval
rateRatio = diffRemote / diffLocal
rateRatioFIR = FIR_filter(rateRatio) // 16-tap mean
PTP_CLOCK_SetDriftPPB((rateRatioFIR - 1.0) * 1e9)
// Clock offset
offset = t2 - t1
// Servo state machine
switch syncStatus:
case UNINIT:
if runs >= 16:
// compute crystal-drift compensation
mac_ti = floor(40.0 * rateRatioFIR)
mac_tisubn = frac(40.0 * rateRatioFIR) * 16_777_216
schedule FOL_ACTION_SET_CLOCK_INC
syncStatus = MATCHFREQ
case MATCHFREQ:
if |offset| > 100_000_000 ns: hardResync = 1
else: syncStatus = HARDSYNC
case HARDSYNC (and beyond):
if |offset| > 0x3FFF_FFFF:
syncStatus = UNINIT // full reset
elif |offset| > 16_777_215 ns:
ta = sign(offset) | min(|offset|, 16_777_215)
schedule FOL_ACTION_ADJUST_OFFSET
elif |offset| > 300 ns:
ta = FIR_coarse_filter(offset)
schedule FOL_ACTION_ADJUST_OFFSET
syncStatus = COARSE
else:
ta = FIR_fine_filter(offset)
schedule FOL_ACTION_ADJUST_OFFSET
syncStatus = FINE
if hardResync:
// Hard-set LAN865x clock directly to GM time
fol_reg_values.tsl = t1.seconds
fol_reg_values.tn = t1.nanoseconds
schedule FOL_ACTION_HARD_SYNC
hardResync = 0
//=======================================================================
// FOLLOWER — register-write serialiser (called every 1 ms)
//=======================================================================
PROCEDURE PTP_FOL_Service():
switch fol_pending_action:
FOL_ACTION_HARD_SYNC:
WriteRegister(MAC_TSL, t1_seconds) // set LAN865x seconds
WriteRegister(MAC_TN, t1_nanosecs) // set LAN865x nanoseconds
FOL_ACTION_SET_CLOCK_INC:
WriteRegister(MAC_TISUBN, tisubn) // sub-nanosecond increment
WriteRegister(MAC_TI, ti) // nanosecond increment
FOL_ACTION_ADJUST_OFFSET:
WriteRegister(MAC_TA, sign_bit | |offset|) // signed offset step
FOL_ACTION_ENABLE_PPS:
WriteRegister(PPSCTL, 0x7D) // enable 1PPS output
//=======================================================================
// SOFTWARE PTP CLOCK (ptp_clock.c)
//=======================================================================
PROCEDURE PTP_CLOCK_Update(wallclock_ns, sys_tick):
s_anchor_wc_ns = wallclock_ns // last known PTP time
s_anchor_tick = sys_tick // TC0 tick captured at that moment
s_valid = true
FUNCTION PTP_CLOCK_GetTime_ns():
delta_tick = SYS_TIME_Counter64Get() - s_anchor_tick
delta_ns = delta_tick * (50 / 3) // 60 MHz → 50/3 ns per tick (exact)
return s_anchor_wc_ns + delta_ns
ptpSync_ct (ptp_fol_task.h:222–228) — follower timestamp store:
typedef struct {
timeStamp_t origin; // t1: GM send time (from FollowUp)
timeStamp_t origin_prev; // t1 from the previous Sync cycle
timeStamp_t receipt; // t2: local receive time (from RTSA)
timeStamp_t receipt_prev; // t2 from the previous Sync cycle
} ptpSync_ct;PTP_RxFrameEntry_t (ptp_ts_ipc.h:33–39) — IPC between driver and app:
typedef struct {
uint8_t data[128]; // raw frame bytes
uint16_t length;
uint64_t rxTimestamp; // RTSA: sec[63:32] | ns[31:0]
uint64_t sysTickAtRx; // TC0 tick captured at SYNC arrival
bool pending; // true = app must process this frame
} PTP_RxFrameEntry_t;syncMsg_t / followUpMsg_t (ptp_fol_task.h:183–198) — PTP wire formats:
typedef struct {
ptpHeader_t header;
ptpTimeStamp_t originTimestamp; // zero in SYNC; filled in FollowUp
} syncMsg_t;
typedef struct {
ptpHeader_t header;
ptpTimeStamp_t preciseOriginTimestamp; // exact t1 from GM hardware clock
tlv_followUp_t tlv; // Organisation-Specific TLV
} followUpMsg_t;GM LAN865x hardware clock (TTSCA registers)
──▶ t1 (seconds + nanoseconds of the SYNC send event)
──▶ FollowUp frame (preciseOriginTimestamp = t1 + 7 650 ns)
──▶ PTP_CLOCK_Update(t1 + anchor_offset, TC0 tick) ← GM software clock
FOL LAN865x hardware clock (RTSA in SPI footer)
──▶ t2 (64-bit: sec[63:32] | ns[31:0])
──▶ g_ptp_raw_rx.rxTimestamp
──▶ handlePtp → processFollowUp
──▶ offset = t2 − t1
──▶ servo → write MAC_TA / MAC_TI / MAC_TSL
──▶ PTP_CLOCK_Update(t2, sysTickAtRx) ← FOL software clock
Software PTP Clock (ptp_clock.c — both boards):
anchor (wc_ns, tick) + TC0 interpolation
──▶ PTP_CLOCK_GetTime_ns() ← used by CLI ptp_time and ptp_time_test.py