Skip to content

Client-mode AAAA lookups return synthetic 'NS localhost.loki.' instead of real answer/SERVFAIL (v0.9.14) #2288

Description

@cpradmin

Client-mode AAAA lookups for arbitrary .loki services return a synthetic NS localhost.loki. record instead of a real answer or SERVFAIL/NXDOMAIN

Summary

On lokinet 0.9.14, client-mode AAAA queries against arbitrary .loki service addresses — both a target service and a separately-tested, independently-documented "known good" control address — consistently return a synthetic authoritative NS localhost.loki. record instead of either a real resolved address or a clean failure (SERVFAIL/NXDOMAIN). The router's own logs show genuine network-level lookup attempts against multiple peer relays failing ("failed to lookup ... from ...snode"), while the client's own service simultaneously connects to the network successfully (introset published, paths built). The bug reproduces identically across two independent, unrelated environments.

Environment 1 — bare metal, static binary

  • OS: Nobara Linux 44 (Fedora-based), kernel 7.1.3
  • lokinet: v0.9.14, official static release binary (lokinet-linux-amd64-v0.9.14.tar.xz)
  • Install: manual, systemd service, client-mode config (lokinet -g -f, no -r)
  • DNS wiring tested two ways: systemd-resolved split-DNS (resolvectl dns/domain → 127.3.2.1), and a dedicated dnsmasq forwarder (server=/loki/127.3.2.1) — both produced the identical result below
  • Also tested via a separate SOCKS5 setup (lokinet + Dante in a container, soren-work/lokinet-proxy) — same underlying failure, different symptom shape (see "SOCKS5" below)

Environment 2 — clean LXC container, official package

  • OS: fresh Ubuntu 24.04 LXC container (Proxmox), no prior configuration
  • lokinet: v0.9.14, installed via the official apt repo (deb.oxen.io), not the static binary
  • DNS: lokinet's own bundled lokinet-resolvconf integration (ran automatically on install)
  • Router status at test time: known/connected: 25/3, paths/endpoints 8/0, multiple paths built in under 1.5s

Reproduction

dig @127.3.2.1 AAAA <any-.loki-service-address>

Result (both environments, both target and control addresses):

;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: ...
;; flags: qr aa rd ra; QUERY: 1, ANSWER: 1, AUTHORITY: 0, ADDITIONAL: 0

;; ANSWER SECTION:
<queried-address>.loki. 1 IN	NS localhost.loki.
  • status: NOERROR with the aa (authoritative) flag set — the resolver is confidently returning this as a final answer, not signaling a pending lookup or failure.
  • Reproduced with two different query tools (dig and host) — identical output, ruling out a client-tool quirk.
  • A plain A query (not AAAA) for the same name correctly returns clean NXDOMAIN, as expected since .loki addresses are IPv6-only — so the resolver does distinguish query types; the bug is specific to the AAAA path.
  • Tested against two separate addresses: the intended target and an independently-sourced "known good" test/control address. Both fail identically, ruling out "the target service is just offline" as the explanation.

Router-level evidence (not just a bad DNS response)

journalctl -u lokinet during a lookup attempt shows real, repeated network-level failures against multiple different peer relays for the same query, e.g.:

[...] endpoint:<our-endpoint>.loki failed to lookup <target>.loki from <relay-1>.snode order=0
[...] endpoint:<our-endpoint>.loki failed to lookup <target>.loki from <relay-2>.snode order=1

This happens while the same client is simultaneously publishing its own introset successfully and building paths, i.e. it is genuinely connected and participating in the network — the failure is specific to looking up other services' introsets, not general connectivity.

SOCKS5 path (separate symptom, same underlying failure)

Testing the same addresses through a SOCKS5 proxy (lokinet + Dante, no DNS involved at the client level — resolution happens proxy-side) fails differently but consistently: curl --socks5-hostname returns SOCKS5 reply code 4 (host unreachable), and the Dante log shows a clean could not resolve hostname "..." : Name or service not known. Same two addresses (target + control), same result.

Ruled out (with evidence, not assumption)

  • DNS client tooling — identical result via dig and host.
  • Local DNS forwarding config — identical result via direct query to 127.3.2.1, systemd-resolved split-DNS, a dedicated dnsmasq forwarder, and the official lokinet-resolvconf integration.
  • Target service being offline — a second, independently-documented control address fails identically.
  • Underlying Oxen network health — checked live via the public block explorer at time of testing: 1,046 active service nodes, full checkpoint consensus (20/20 signatures), normal block production cadence. Not a network-wide outage.
  • Host firewall — inspected the relevant DOCKER-USER/forward chains directly; 0 packets matched on any rule that could plausibly affect this traffic.
  • NetBird (a separate overlay/VPN client running on the host) — fully stopped and retested; identical result with it down.
  • MTU/fragmentation — path MTU to an actual bootstrap peer confirmed 1500 bytes end-to-end over 13 hops; both tun interfaces correctly at 1500.
  • Environment-specific interference — reproduced from scratch on a completely independent machine, OS, install method, and DNS integration path (see Environment 2).

Suspected root cause

The literal template string %s 10800 IN NS localhost. is present in the compiled lokinet binary (strings /usr/local/bin/lokinet), alongside the standard localhost./127.in-addr.arpa reverse-DNS boilerplate normally used only for local loopback responses. This template's shape (<name> <ttl> IN NS localhost.) matches the malformed answer exactly. It appears this synthetic local/loopback response template is being incorrectly returned for real service-lookup queries under some condition (possibly related to lookup timeout/failure handling on the AAAA path specifically, given A queries return a clean NXDOMAIN in the same conditions).

Impact

Client-mode .loki service resolution via the standard DNS interface (127.3.2.1) appears to be non-functional for arbitrary service addresses in v0.9.14, both via direct/forwarded DNS and via the SOCKS5 proxy path, across at least two independent, differently-configured environments.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions