disable_local_auth: true, the one Azure OpenAI setting that separates a demo from a production deployment

When we scoped the attack generation service for Noble Lynx, we set one rule before writing a line of code, no API keys anywhere in the system, full stop. That rule sounds obvious on paper. In practice it forces a series of decisions that most Azure OpenAI deployments never make, and the gap between those two states is exactly what separates a working demo from something you can put in front of an enterprise security team.
Here is the setting, why it matters more than almost anything else in your Azure OpenAI configuration, and the exact migration path if your current deployment still accepts a key.
The setting
Every Azure OpenAI resource has a property called disableLocalAuth. In the portal it sits under Networking. In Bicep it is one line.
resource openai 'Microsoft.CognitiveServices/accounts@2024-10-01' = {
name: 'servicename-openai-weu'
location: 'westeurope'
kind: 'OpenAI'
sku: {
name: 'S0'
}
properties: {
disableLocalAuth: true
publicNetworkAccess: 'Disabled'
customSubDomainName: 'servicename-openai-weu'
}
}
disableLocalAuth false, the default on every new resource, means the endpoint accepts API keys. Anyone holding the key can call the model. No identity behind the call, no per service audit trail, no way to revoke access for one consumer without rotating the key for all of them.
disableLocalAuth true removes that path entirely. The resource only accepts Microsoft Entra ID tokens. A valid key, if one still exists, simply stops working. Authentication becomes proving you are an identity that has been explicitly granted a role, not proving you know a string that was generated once and never expires on its own.
This is the highest leverage change available on an Azure OpenAI deployment, and it is also the one most teams skip, because the key based version is faster to wire up during the first sprint.
Why we treated this as non negotiable
Noble Lynx audits other companies’ AI deployments for exactly this class of gap, unscoped credentials, standing access nobody remembers granting, no attribution when something goes wrong. Shipping our own attack generation service with a long lived API key sitting in a secret store would have meant our own product violated the first finding in our own report template. That is not a brand risk we were willing to carry, so the architecture decision was made before any code existed.
The attack generation agent calls Azure OpenAI dozens of times per scan, crafting and scoring prompt injection variants. Every one of those calls needed to be attributable to that specific agent, in that specific environment, with permissions that could be revoked in isolation if the agent ever needed to be pulled out of rotation mid investigation.
What breaks when you flip it, and the migration
If existing code authenticates with an API key, flipping disableLocalAuth to true breaks every call immediately. There is no fallback. That is the point, but it means this is a small migration, not a config toggle you flip and walk away from.
Before, key based auth
import os
from openai import AzureOpenAI
client = AzureOpenAI(
api_key=os.environ["AZURE_OPENAI_KEY"],
api_version="2024-10-21",
azure_endpoint="<https://servicename-openai-weu.openai.azure.com/>"
)
response = client.chat.completions.create(
model="gpt-4.1",
messages=[{"role": "user", "content": prompt}]
)
Every place this key exists, a secret manager, a CI pipeline, a developer’s local shell, is a place it can leak. And because the key has no built in scoping, a leak means unrestricted inference against the whole resource.
After, Managed Identity
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
from openai import AzureOpenAI
credential = DefaultAzureCredential()
token_provider = get_bearer_token_provider(
credential,
"<https://cognitiveservices.azure.com/.default>"
)
client = AzureOpenAI(
azure_ad_token_provider=token_provider,
api_version="2024-10-21",
azure_endpoint="<https://servicename-openai-weu.openai.azure.com/>"
)
response = client.chat.completions.create(
model="gpt-4.1",
messages=[{"role": "user", "content": prompt}]
)
No secret in the code, no secret in the environment. DefaultAzureCredential resolves to whatever identity the process is actually running as, a Managed Identity in Azure, an az login session for local development. The resulting token is short lived, automatically refreshed, and scoped to the resource it was issued for.
The step that actually determines whether this works
The code change is half the migration. Nothing succeeds until the calling identity has a role assignment on the target resource, and this is where most first attempts produce a wall of 403s.
az role assignment create \
--assignee-object-id <managed-identity-principal-id> \
--assignee-principal-type ServicePrincipal \
--role "Cognitive Services OpenAI User" \
--scope "/subscriptions/<sub-id>/resourceGroups/<rg>/providers/Microsoft.CognitiveServices/accounts/servicename-openai-weu"
Cognitive Services OpenAI User is the minimum role for inference, it cannot modify the deployment, change network rules, or touch other resources in the group. For the attack generation agent specifically, that role on that one resource is the entire grant. It has no path to the Cosmos DB instance holding scan state, no path to the regulatory mapping service’s deployment. If that identity were ever compromised mid scan, the blast radius is one model endpoint, not the subscription.
Local development without bringing the key back
The usual pushback is that this makes local development worse. It does not, if scoped correctly.
az login
DefaultAzureCredential checks for an active CLI session as one of its fallback paths. Local development authenticates as the developer, with whatever roles that developer holds, not as a shared secret every contributor also has sitting in a dotenv file. Offboarding becomes disabling an account, not rotating a key across every service that referenced it.
What this changes in practice
Three outcomes mattered more than expected once this was in place.
Attribution. Every call now resolves to a specific Managed Identity in diagnostic logs, which maps to a specific agent in a specific environment. Debugging an unexpected scan result starts with knowing exactly which service made which call.
Revocability. Pulling one agent’s access means removing one role assignment, not coordinating a key rotation across every consumer that shared it.
An answer that holds up under client scrutiny. Noble Lynx sells audits to companies who will ask how we secure our own infrastructure. No API keys, every service authenticates through Managed Identity with least privilege role assignments, scoped per resource, is a complete answer. A shared key in a vault invites a longer conversation about who can read that vault.
The checklist
For any Azure OpenAI deployment touching real data, in order.
- Set disableLocalAuth true on every Cognitive Services resource.
- Replace every api_key call site with DefaultAzureCredential and get_bearer_token_provider.
- Grant Cognitive Services OpenAI User, not Contributor, not Owner, scoped to the specific resource, not the resource group.
- Confirm az login covers local development before removing the old key from wherever it lives.
- Delete, not just stop referencing, any existing key once nothing depends on it.
- Confirm diagnostic settings capture the Audit log category, so the identity to call mapping this enables is actually queryable later.
None of this is exotic. It is one resource property and one credential object away from where most deployments already sit. The gap between demo and production on Azure OpenAI is almost always this setting, not a missing product, not a complex control, a default nobody flipped because the key worked fine in the sprint where shipping fast mattered more.