During a review of the AI Foundry Landing Zone and AI Gateway Landing Zone architectures , we identified several areas where the current documentation can lead to different interpretations by field teams and customers.
Key Feedback:
Clarify when AI Gateway should be considered mandatory versus optional.
Provide clear guidance for centralized vs decentralized AI operating models.
Document recommended placement of AI Gateway, Foundry resources, and knowledge bases.
Clarify networking options (Hub-Spoke via Firewall vs Direct VNet Peering) along with governance, performance, and latency trade-offs.
Align AI Landing Zone diagrams with Enterprise Scale/ALZ guidance, particularly around WAF and Application Gateway placement.
Add architecture decision trees or reference scenarios to help field teams consistently recommend the correct deployment model.
Improve documentation structure so all deployment options and recommendations are easier to discover and understand.
The architecture does not clearly define:
Model A: Centralized AI Platform
Central AI Gateway
Central Foundry
Central model hosting
Shared access for application teams
Model B: Distributed AI Platform
Individual Foundry deployments per application/business unit
Centralized AI Gateway governance
Application-specific knowledge bases
The guidance should explicitly describe when each model should be used.
Current patterns show traffic flowing through connectivity subscriptions and firewall inspection paths, but the documentation does not clearly explain:
Direct VNet peering options
Hub-spoke routing alternatives
East-West traffic considerations
Performance implications
A typical request may follow:
Application → AI Gateway → Foundry → AI Gateway → Application
In firewall-inspected environments, the number of routing hops increases significantly.
Suggested Improvement
Document:
Latency implications
Cost implications
Security implications
Recommended patterns for:
BFSI customers
Manufacturing customers
General enterprise customers
Currently there is no guidance on:
Option A
Hub-and-spoke with Azure Firewall inspection
Pros:
Maximum governance
East-West inspection
Security compliance
Cons:
Increased latency
Increased RTTs
Additional firewall cost
Option B
Direct VNet Peering
Pros:
Lower latency
Lower operational cost
Simpler routing
Cons:
Reduced inspection
A documented decision framework would help field teams recommend architectures consistently.
The documentation should clarify whether AI Gateway should be deployed:
In a dedicated AI governance subscription
In the hub/connectivity subscription
In application subscriptions
The reasoning and trade-offs for each option should be documented.
The current diagrams show WAF/Application Gateway placement within application subscriptions.
This appears inconsistent with common Azure Landing Zone designs where ingress services are typically placed within connectivity/hub subscriptions.
The documentation should either:
Align with existing ALZ guidance, or
Clearly explain why AI Landing Zone guidance intentionally differs.
Expected Outcome: Provide prescriptive architecture guidance, deployment recommendations, and networking patterns that allow customers and Microsoft field teams to consistently implement AI Foundry and AI Gateway Landing Zones while balancing governance, security, performance, and operational complexity.
Yusuf (@whereisyusuf)
During a review of the AI Foundry Landing Zone and AI Gateway Landing Zone architectures , we identified several areas where the current documentation can lead to different interpretations by field teams and customers.
Key Feedback:
Clarify when AI Gateway should be considered mandatory versus optional.
Provide clear guidance for centralized vs decentralized AI operating models.
Document recommended placement of AI Gateway, Foundry resources, and knowledge bases.
Clarify networking options (Hub-Spoke via Firewall vs Direct VNet Peering) along with governance, performance, and latency trade-offs.
Align AI Landing Zone diagrams with Enterprise Scale/ALZ guidance, particularly around WAF and Application Gateway placement.
Add architecture decision trees or reference scenarios to help field teams consistently recommend the correct deployment model.
Improve documentation structure so all deployment options and recommendations are easier to discover and understand.
The architecture does not clearly define:
Model A: Centralized AI Platform
Central AI Gateway
Central Foundry
Central model hosting
Shared access for application teams
Model B: Distributed AI Platform
Individual Foundry deployments per application/business unit
Centralized AI Gateway governance
Application-specific knowledge bases
The guidance should explicitly describe when each model should be used.
Current patterns show traffic flowing through connectivity subscriptions and firewall inspection paths, but the documentation does not clearly explain:
Direct VNet peering options
Hub-spoke routing alternatives
East-West traffic considerations
Performance implications
A typical request may follow:
Application → AI Gateway → Foundry → AI Gateway → Application
In firewall-inspected environments, the number of routing hops increases significantly.
Suggested Improvement
Document:
Latency implications
Cost implications
Security implications
Recommended patterns for:
BFSI customers
Manufacturing customers
General enterprise customers
Currently there is no guidance on:
Option A
Hub-and-spoke with Azure Firewall inspection
Pros:
Maximum governance
East-West inspection
Security compliance
Cons:
Increased latency
Increased RTTs
Additional firewall cost
Option B
Direct VNet Peering
Pros:
Lower latency
Lower operational cost
Simpler routing
Cons:
Reduced inspection
A documented decision framework would help field teams recommend architectures consistently.
The documentation should clarify whether AI Gateway should be deployed:
In a dedicated AI governance subscription
In the hub/connectivity subscription
In application subscriptions
The reasoning and trade-offs for each option should be documented.
The current diagrams show WAF/Application Gateway placement within application subscriptions.
This appears inconsistent with common Azure Landing Zone designs where ingress services are typically placed within connectivity/hub subscriptions.
The documentation should either:
Align with existing ALZ guidance, or
Clearly explain why AI Landing Zone guidance intentionally differs.
Expected Outcome: Provide prescriptive architecture guidance, deployment recommendations, and networking patterns that allow customers and Microsoft field teams to consistently implement AI Foundry and AI Gateway Landing Zones while balancing governance, security, performance, and operational complexity.
Yusuf (@whereisyusuf)