Skip to content

Clarification Required for AI Landing Zone & AI Gateway Architecture Guidance #137

Description

@faizaan03

During a review of the AI Foundry Landing Zone and AI Gateway Landing Zone architectures , we identified several areas where the current documentation can lead to different interpretations by field teams and customers.

Key Feedback:

Clarify when AI Gateway should be considered mandatory versus optional.
Provide clear guidance for centralized vs decentralized AI operating models.
Document recommended placement of AI Gateway, Foundry resources, and knowledge bases.
Clarify networking options (Hub-Spoke via Firewall vs Direct VNet Peering) along with governance, performance, and latency trade-offs.
Align AI Landing Zone diagrams with Enterprise Scale/ALZ guidance, particularly around WAF and Application Gateway placement.
Add architecture decision trees or reference scenarios to help field teams consistently recommend the correct deployment model.
Improve documentation structure so all deployment options and recommendations are easier to discover and understand.

The architecture does not clearly define:

Model A: Centralized AI Platform

Central AI Gateway
Central Foundry
Central model hosting
Shared access for application teams

Model B: Distributed AI Platform

Individual Foundry deployments per application/business unit
Centralized AI Gateway governance
Application-specific knowledge bases

The guidance should explicitly describe when each model should be used.

Current patterns show traffic flowing through connectivity subscriptions and firewall inspection paths, but the documentation does not clearly explain:

Direct VNet peering options
Hub-spoke routing alternatives
East-West traffic considerations
Performance implications

A typical request may follow:

Application → AI Gateway → Foundry → AI Gateway → Application

In firewall-inspected environments, the number of routing hops increases significantly.

Suggested Improvement

Document:

Latency implications
Cost implications
Security implications
Recommended patterns for:
BFSI customers
Manufacturing customers
General enterprise customers

Currently there is no guidance on:

Option A

Hub-and-spoke with Azure Firewall inspection

Pros:

Maximum governance
East-West inspection
Security compliance

Cons:

Increased latency
Increased RTTs
Additional firewall cost
Option B

Direct VNet Peering

Pros:

Lower latency
Lower operational cost
Simpler routing

Cons:

Reduced inspection

A documented decision framework would help field teams recommend architectures consistently.

The documentation should clarify whether AI Gateway should be deployed:

In a dedicated AI governance subscription
In the hub/connectivity subscription
In application subscriptions

The reasoning and trade-offs for each option should be documented.

The current diagrams show WAF/Application Gateway placement within application subscriptions.

This appears inconsistent with common Azure Landing Zone designs where ingress services are typically placed within connectivity/hub subscriptions.

The documentation should either:

Align with existing ALZ guidance, or
Clearly explain why AI Landing Zone guidance intentionally differs.

Expected Outcome: Provide prescriptive architecture guidance, deployment recommendations, and networking patterns that allow customers and Microsoft field teams to consistently implement AI Foundry and AI Gateway Landing Zones while balancing governance, security, performance, and operational complexity.

Yusuf (@whereisyusuf)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions