Your Azure OpenAI private endpoint may still be using the public internet
By Matthew Gray on Oct 6, 2026
A regional healthcare group had Azure OpenAI private endpoints in both development and production, and its architecture diagrams showed AI traffic staying inside the virtual network. In practice, every API call from those networks was going to a public IP address. A private endpoint only makes traffic private if DNS sends clients to it, and in this estate two separate DNS mistakes meant it did not.
Why “the private endpoint exists” proves very little
A private endpoint is a network interface with a private IP address in your subnet. Clients don’t connect to that IP directly. They connect to the service’s normal name, such as <account>.openai.azure.com, and depend on DNS to return the private address.
Public DNS for that name returns a CNAME to <account>.privatelink.openai.azure.com. Inside your network, a private DNS zone called privatelink.openai.azure.com is meant to answer that second name with the endpoint’s private IP. If the zone has no matching record, or your resolver never asks Azure about the zone, resolution carries on down the public CNAME chain and returns a public IP.
Usually nothing breaks when this happens. In this case the OpenAI resources still had public network access enabled, with a VNet service-endpoint rule that allowed the application subnets. The calls succeeded, the application worked, and nobody had reason to look. The only symptom was a DNS answer nobody was checking.
Misconfiguration 1: records named after the endpoint, not the account
The private DNS zone had an A record for each environment, but each record was named after the private endpoint resource (the pe-… name) instead of the OpenAI account. Clients never query the endpoint’s resource name. They query the account name. So the records sat in the zone, pointing at the correct IPs, and no client lookup ever reached them.
Here is how it happened. Neither endpoint had a working privateDnsZoneGroup. Development had none at all. Production had a zone group with an empty zone list, which looks fine at a glance and does nothing. Without a zone group, Azure does not create or maintain the A record, so someone added the records by hand and used the wrong name.
This failure is easy to repeat anywhere records are managed by hand. The record has to match the target resource’s name, and it has to be updated if the endpoint’s IP changes.
Misconfiguration 2: the resolver never asked Azure
Fixing the record name was not enough. The VNets used custom DNS servers: domain controllers in the hub, reached over peering. Those servers had no conditional forwarder for the privatelink zone, so they resolved the name recursively from public root hints, followed the public CNAME chain, and returned a public IP.
The private DNS zone is only visible through Azure’s platform resolver at 168.63.129.16. If your custom DNS doesn’t forward the relevant zones to it, and the VNet’s DNS settings don’t include it, your private zone might as well not exist.
Linux VMs added one more wrinkle. systemd-resolved queries servers in order and accepts the first valid answer. Because the domain controllers returned a valid (public) answer, it never tried any other server. Our first per-VM routing rule matched ~privatelink.openai.azure.com, and it did not work. Routing is decided by the name in the original query, which ends in .openai.azure.com, not by the CNAME targets that the upstream resolver follows on its own. Widening the match to ~openai.azure.com fixed it.
How to check yours
Run the lookup from inside the VNet, on a VM or container that actually calls the service. A lookup from your laptop tells you nothing about what your workloads see.
# Linux
dig +short <account>.openai.azure.com
getent hosts <account>.openai.azure.com
# Windows
Resolve-DnsName <account>.openai.azure.com
nslookup <account>.openai.azure.com
A correct answer is short: a CNAME to <account>.privatelink.openai.azure.com, then an address inside your VNet’s private range. A broken answer usually runs on past the privatelink name to regional API hostnames and traffic-manager names, and ends at a public IP. If you see that, work through three causes in order:
- Endpoint connection state. Check that the private endpoint connection is approved.
- Zone record name. List the A records in the
privatelinkzone and confirm each one is named after the account, not the endpoint. Check whether each endpoint has a zone group, and whether that group actually lists a zone. - Resolver chain. Find out which DNS servers the VNet hands out, and whether they forward
privatelink.openai.azure.com(oropenai.azure.com) to168.63.129.16. Confirm the zone is linked to the VNet the client is in.
While you are there, run the same check on every other private endpoint. This is a pattern, not something specific to OpenAI. Key Vault and Storage endpoints rely on the same DNS chain and fail the same way.
The fix
We made the changes in development, verified them, and then repeated them in production:
- Attach a
privateDnsZoneGroupto each private endpoint that references theprivatelink.openai.azure.comzone. Azure then creates the correctly named record and keeps it in step with the endpoint’s IP. In production, we had to delete the empty zone group before creating a working one. - Delete the misnamed manual records so the zone contains only records clients actually query.
- Append
168.63.129.16to the VNet’s DNS server list. The client’s policy was append-only: internal DNS stayed first. - Route only the OpenAI names to Azure DNS on each VM, with a narrow
systemd-resolveddrop-in. All other lookups still go to the domain controllers first.
[Resolve]
DNS=168.63.129.16
Domains=~openai.azure.com
After the change, the account name resolved to the endpoint’s private IP in both environments, TCP 443 to that private IP succeeded, and unrelated names still resolved through the domain controllers.
The per-VM drop-in is a stopgap, and we said so in the change record. The durable fixes are conditional forwarders on the internal DNS servers, or an Azure DNS Private Resolver, so new VMs work without per-host configuration. After that, add an Azure Policy guardrail that requires a zone group on every new private endpoint. Once every legitimate client is confirmed on the private path, disable public network access on the resource so the public path is closed for good.
Zero downtime, under change control
The only thing these changes alter is which IP a name resolves to. The service was already reachable on both paths, so moving clients from the public path to the private one did not interrupt service. We still treated it as a production change: a formal change notice, development before production, a before-and-after record of each step, and a written rollback for every one. That record is what lets an auditor confirm the fix happened, rather than taking our word for it.
How we can help
If your architecture says your AI or data traffic is private, the DNS answer from inside the network is the evidence that settles it. We do this kind of read-only check across a whole subscription’s private endpoints, then fix what we find under your change process. You can read how this fit into a wider audit-readiness engagement in our healthcare SIEM case study, or get in touch.