Serving must not depend on the Manager
The query path stays on the DNS node and its last verified state.

AUTHORITATIVE DNS / GLOBAL TRAFFIC
API-first management of DNS desired state, validation, immutable revisions, blue/green delivery, health-aware routing, DNSSEC and node lifecycle in a single process, while the query path remains independent of the Manager.
4SO GeoDNS components: a query path independent from the control path
Product problem
Authoritative DNS must be fast and always available, while its changes must be precise, auditable and reversible. GeoDNS separates the configuration path from the serving path.
The query path stays on the DNS node and its last verified state.
Records and policies first become a revision, then are staged and verified.
Endpoint health and eligibility come from accepted observations, not from simulated browser state.
Keys, signing, rollover and recovery must be as controlled as zone management.
Capability snapshot
Beyond authoritative records, GeoDNS treats delivery, traffic management and node lifecycle as part of the same product.
Forward/reverse zones, typed RRsets and controlled import/export.
Changesets, snapshot hashes and revisions for audit and rollback.
Stage on the inactive color, read back state, then cut over atomically.
Endpoints, pools, regions, health perspectives and answer policy.
Signing, DS evidence, rollover and restore/reconcile.
TSIG, AXFR/IXFR, catalog zones and secondary provider lifecycle.
Onboard, drain, replace, remove and fleet rollout.
Health, metrics, alerts, incidents and synthetic DNS probes.
The normal query path stays short: Resolver → DNS node → active runtime → local state. No regular query depends on the Manager or a cross-region database to be answered.
Deployment model
The minimum product topology is simple, and HA or multi-region is not a prerequisite; nodes and regions are added incrementally.
| Profile | Control Plane | Serving Plane | Use case |
|---|---|---|---|
| Basic | 1 Manager | 1 DNS node | Simple, complete authoritative DNS setup |
| Multi-node | Central Manager | Multiple DNS nodes | Serving redundancy and controlled rollout |
| Multiple sites | Single control plane | Nodes across multiple sites | Failure-domain separation and proximity to resolvers |
| Multi-region | Central management with regional processes | Independent serving in multiple regions | Global traffic and broad resilience |
Architecture
A DNS change must pass through validation, revision and delivery. DNS answers, however, are served directly from the node's verified runtime.
| Layer | Component | Responsibility | Fail-safe behavior |
|---|---|---|---|
| Product API | FastAPI + scoped RBAC | Zone/RRset, traffic policy, node lifecycle and typed validation | The browser never mutates the DNS runtime directly |
| Authority | PostgreSQL + transactional outbox | Desired state, audit, job/revision state | Mutation and event publication within one transactional boundary |
| Delivery | Worker + mTLS Node Agent | Stage inactive color, read-back, verify and atomic cutover | Failed delivery stops before serving |
| DNS Edge | dnsdist | Stable serving entry and selection of the active PowerDNS color | The query path does not depend on the Manager |
| Authoritative Runtime | PowerDNS blue/green + local PostgreSQL | Serves the last verified revision | A control-plane outage keeps DNS answering from the last healthy state |
| Security | DNSSEC · TSIG · transfer policy | Signing, rollover and zone-transfer trust | Secret material never enters Panel/Audit/Evidence |
Geo-routing & Health Contract. DNS decisions are read from the compiled policy and a local snapshot; the query path never runs a network health check at answer time.
| Domain | Input / Mechanism | Technical contract |
|---|---|---|
| Client locality | ECS or Recursive Resolver IP | Valid ECS is used when permitted; otherwise the resolver IP is the basis for selection. |
| Geo / network match | CIDR, ASN/Organization, Country/Subdivision/Continent/City | Precise overrides and geography are evaluated by numeric priority in deterministic order. |
| Native GeoIP | PowerDNS geoip backend + Country/ASN MMDB | GeoIP sits alongside gpgsql; zone data is not moved from the primary authority into a YAML/GeoIP backend. |
| Health | TCP/UDP/HTTP/HTTPS/TLS/DNS/gRPC/ICMP probes | Health is computed outside the query path, and states such as HEALTHY/DEGRADED/SUSPECT/UNHEALTHY are recorded in a versioned snapshot. |
| Policy compiler | Structured policy → trusted runtime material | Arbitrary Lua is not accepted; the compiler produces reproducible output with a hash bound to the revision. |
| Simulation | Resolver/ECS/time/health/dataset inputs | Rule match, selected pool/endpoint, fallback path, answer and TTL can be simulated before publish. |
| ECS cache isolation | Policy-aware cache keying | Answers that depend on the client network must not leak through the cache across unrelated ECS contexts. |
Lifecycle
Going from edit to publish is more than a few clicks; validation, revision, stage, state read-back and recovery each have a defined state.
Capability profiles
Every layer, from authoritative data to global traffic and security, has its own operational profile.
Zone and RRset management with revisions.
Answers based on endpoint and policy.
Trust and transfer security inside the product.
Lifecycle of serving nodes.
Safe, reversible publishing.
Protecting management state.
Job, cutover and recovery flow
Mutations are idempotent and project-scoped; node mutation, fleet rollout and certificate rotation never stay open concurrently on the same node. Worker ownership is fenced with a lease and an execution epoch.
DNS read-back happens before the active color is switched.
The query path never runs a network probe at request time.
DR actions are bound to the readiness hash and topology freeze.
The change is staged on the inactive color first.
Geo/network policy is simulated before publish.
A node joins the fleet only after trust and serving verification.
Key lifecycle runs with overlap and serving acknowledgement.
Restore and site failover are bound to readiness evidence.
Health and node replacement directly affect serving quality and need their own flows.
Health is collected outside the query hot path and turned into a versioned snapshot.
The old node leaves only after the replacement is ready and serving is verified.
| Action / Flow | Admission / Preconditions | Execution / Lock | Success Criterion | Failure / Recovery |
|---|---|---|---|---|
| Publish Zone / RRset | Typed validation, revision and policy compile before delivery | Stage on inactive color → node read-back → atomic cutover | Authoritative answer/TTL and active revision confirmed from the serving path | A failed stage does not change serving; rollback to the preserved healthy color |
| Traffic Policy | Structured policy, priority and simulation without conflict | Compiler material is versioned; health snapshot is consumed outside the query path | Matched rule/pool/endpoint and synthetic DNS behavior match the revision | Arbitrary Lua is rejected; policy conflict fails before publish |
| Node Onboarding / Replace | Identity/mTLS, preflight and node mutation availability | Durable job with node lock; onboarding/day-2/fleet/cert-rotation overlap is rejected | Node observed healthy and rollout eligibility clear | Node onboarding failure does not become a generic retry; a fresh bootstrap may be required |
| DNSSEC / Certificate | Key/identity lifecycle and approval for sensitive operations | Stage/overlap/ack and owner-specific rotation workflow | Signing/serving or certificate observation matches the expected generation | Stale rotation/concurrent fleet mutation rejected; secret material removed from the public job projection |
| Control-plane Restore | Backup must be READY, repository-verified and carry snapshot/hash evidence | Preflight → approval → execution → cutover | RESTORED only after preflight binding and runtime verification | Worker interruption/unknown external state → RECONCILE_REQUIRED; blind retry prohibited |
| DR Failover / Failback | Readiness job, topology freeze, fencing, rollback and RPO acknowledgement required | The action is bound to readiness state/hash and profile revision | Active site/traffic/database state reaches the expected phase | Stale readiness or topology mismatch rejected; failback allowed only from FAILED_OVER |
Retry classification is explicit to the operator: SAFE_TO_RETRY, RECONCILE_REQUIRED, FRESH_BOOTSTRAP_REQUIRED or FRESH_HEALTH_REQUEST_REQUIRED. Lease loss in disruptive workflows is never permission for blind replay.
Operations catalog
Operators use the same product process and audit trail to change records, roll out nodes or manage DNSSEC.
4SO GeoDNS controls changes without putting the Control Plane in the DNS answer path.