Academy/VCAP — VCF Architect (3V0-12.26)/Lab: VCF Consumption & Automation Architecture Design
This lab targets VCF 9.0

Lab: VCF Consumption & Automation Architecture Design

VCF 9.0Advancedvcap-advancedvcdx⏱ 120 min

Consumption strategy lab — VCF Automation Cloud Templates, service catalog design, tenant onboarding, governance policies, IaC integration. VCF 9.0.

Objectives

  • Design a VCF Automation service catalog with T-shirt sized VM templates and approval workflows
  • Architect tenant onboarding workflows including project creation, quota assignment, and role mapping
  • Design Cloud Template (YAML) blueprints for standardized VM provisioning with network and storage policy integration
  • Implement governance policies: resource quotas, lease policies, naming conventions, and cost transparency
  • Integrate VCF Automation with NSX VPC for self-service network provisioning

Prerequisites

VCF 9.0 with VCF Automation deployed and integrated with vCenter and NSX. At least one workload domain available for provisioning. NSX VPC (optional — for advanced self-service networking).

Prior labs: vcap-architect-01

Required skills:

  • VCF Automation concepts (Cloud Templates, Service Catalog, Projects)
  • YAML syntax for infrastructure-as-code templates
  • NSX networking basics (segments, VPC)
  • Multi-tenancy and RBAC concepts

Lab Environment

VCF 9.0 management domain with VCF Automation deployed. 1 workload domain with available capacity for VM provisioning. NSX with T0/T1 operational.

Tasks

Task 1 Design Consumption Architecture with VCF Automation

Consumption architecture is where VCF delivers cloud-like agility. The design must balance self-service speed (developers get VMs in minutes) with governance (quotas, naming, cost control). Over-governing kills adoption; under-governing creates sprawl. The right balance depends on organizational maturity.

Design a self-service infrastructure consumption model using VCF Automation — from service catalog design through tenant onboarding to governance enforcement. This transforms VCF from admin-provisioned infrastructure into a cloud-like self-service platform.

Step 1

Design T-shirt sized VM templates. Define standard VM sizes that map to common workload profiles: Small (2 vCPU, 4GB RAM, 60GB disk) — development, testing, lightweight services. Medium (4 vCPU, 8GB RAM, 100GB disk) — application servers, web servers. Large (8 vCPU, 16GB RAM, 200GB disk) — database servers, middleware. XLarge (16 vCPU, 32GB RAM, 500GB disk) — heavy databases, analytics. For each size, define: vSAN storage policy (Bronze/Silver/Gold mapping), network segment assignment (dev/staging/production), OS template (RHEL 9, Windows 2022, Ubuntu 22.04). Document rationale for each size — trace to common workload profiles from capacity planning.

Step 2
Create VCF Automation Cloud Template (YAML). Build a parameterized template: inputs section with dropdowns for size (S/M/L/XL), environment (dev/staging/production), OS (RHEL/Windows/Ubuntu). Resources section: Cloud.vSphere.Machine with flavor mapping to T-shirt size, image mapping to OS template, Cloud.NSX.Network for network assignment based on environment input. Constraints section: tag-based placement to correct workload domain (environment=production → Clinical domain, environment=dev → Development domain). Add custom naming: ${input.environment}-${input.application}-${resource.count.index}. Test: deploy one VM from the template, verify naming, network, and storage policy.
Step 3

Design the service catalog structure. Catalog Organization: (a) Category: 'Virtual Machines' — items: Linux VM (S/M/L/XL), Windows VM (S/M/L/XL), Custom VM (manual sizing). (b) Category: 'Networking' — items: VPC Request (for teams needing isolated network), Load Balancer Request, VPN Connection Request. (c) Category: 'Kubernetes' — items: VKS Cluster (dev/production profiles). (d) Category: 'Day-2 Operations' — items: Resize VM, Add Disk, Snapshot, Decommission. For each catalog item: set icon, description, input form layout, estimated provisioning time. Design Decision: limit catalog to standardized offerings — custom requests go through ticket-based approval to prevent non-standard configurations.

Step 4
Design tenant onboarding workflow. When a new team (tenant) is onboarded: (a) Create VCF Automation Project: Name=team name, administrators=team leads, members=developers. (b) Assign Cloud Zones: map project to workload domain cloud zones (dev → Development domain, production → Clinical domain with approval). (c) Set Resource Quotas: per-project limits (e.g., 'Engineering-Dev': max 50 VMs, 200 vCPU, 800GB RAM, 5TB storage). (d) Configure Approval Policies: Small/Medium VMs = auto-approved; Large/XLarge = team lead approval; Production = infrastructure admin approval. (e) Assign Naming Policy: enforce naming convention per project. (f) Enable Cost Showback: configure cost allocation tags for per-project cost tracking. Document this as a repeatable onboarding checklist.
Step 5

Implement resource governance. (a) Lease Policies: Development VMs expire after 30 days (with 7-day warning email). Staging VMs expire after 90 days. Production VMs have no lease (permanent). (b) Quota Enforcement: when a project hits 80% of quota, notify project admin. At 100%, block new deployments until resources are released. (c) Naming Convention: enforce via VCF Automation custom naming — all VMs must follow: {env}-{team}-{app}-{index} (e.g., dev-alpha-web-001). Reject deployments with non-compliant names. (d) Tag Enforcement: require tags for 'cost-center', 'application-owner', 'data-classification' on every deployment. Configure input forms to capture these as mandatory fields.

Step 6
Design approval workflows. Multi-level approval chain: (a) Level 1 (auto-approve): Small/Medium VMs in development environment → immediate provisioning. (b) Level 2 (team lead): Large/XLarge VMs or staging environment → team lead approves via email notification with 24h SLA. (c) Level 3 (infra admin): any production deployment or custom sizing → infrastructure admin reviews capacity impact and approves. (d) Level 4 (security review): any deployment tagged 'data-classification=PHI' → security team verifies HIPAA controls before approval. Configure approval policies in VCF Automation: Policies → Approval Policies → Add. Set conditions (size, environment, data classification) and approver groups.
Step 7

Integrate with NSX VPC for self-service networking. Design the VPC consumption model: (a) Infrastructure admin pre-provisions: NSX Projects (one per team), IP blocks (172.16.0.0/16 partitioned per project), T0 gateway binding, Edge cluster resources. (b) Tenant self-service: within their NSX Project, tenant admins create VPCs with subnets, NAT rules, and security policies — no ticket to infrastructure team. (c) VCF Automation integration: Cloud Template includes Cloud.NSX.VPC resource — tenant requests a VPC from the catalog, VCF Automation creates the VPC with pre-defined subnet sizes and security policies. (d) Guardrails: maximum 5 VPCs per project, maximum 10 subnets per VPC, IP block allocation monitored for exhaustion.

Step 8

Design cost transparency and showback. (a) Cost Model: define rates per resource unit — $0.02/vCPU/hour, $0.005/GB RAM/hour, $0.01/GB storage/month. These rates should reflect actual infrastructure cost (hardware amortization + licensing + power + cooling + operations, divided by total resource capacity). (b) VCF Operations Integration: create super metrics for per-VM cost calculation (as in VCAP Ops lab). Map costs to VCF Automation projects via tags. (c) Showback Dashboard: per-project monthly cost, top-10 most expensive VMs, cost trend over 6 months, rightsizing savings potential. (d) Chargeback (optional): if the organization does internal billing, configure monthly cost reports per cost center with export to finance system.

Step 9

Design Day-2 operations catalog. After initial provisioning, tenants need ongoing operations: (a) Resize VM: catalog action that changes vCPU/RAM (within quota limits). Requires VM restart warning. (b) Add Disk: catalog action that adds additional VMDK. Maximum 4 additional disks per VM. (c) Snapshot: create/revert/delete snapshots. Limit: max 3 snapshots per VM, max 72h retention (snapshots older than 72h auto-deleted via policy). (d) Decommission: request VM deletion with approval workflow. Backup taken before deletion (7-day retention). (e) Clone to Dev: clone a production VM to development environment with sanitized data. Security review required. Each Day-2 action is a VCF Automation resource action with ABX extensibility for custom logic.

Step 10

Validate the consumption architecture end-to-end. Test scenario: (a) Onboard a new project 'Team-Beta'. (b) Assign quotas: 20 VMs, 80 vCPU, 320GB RAM. (c) Deploy a Medium Linux VM from the catalog — verify auto-approval, naming convention, tag assignment, network placement. (d) Deploy a Large VM — verify team lead approval workflow triggers. (e) Check quota consumption after deployments. (f) Run showback report — verify per-project cost calculation. (g) Execute Day-2 resize — verify quota updated. (h) Let a dev VM expire via lease policy — verify warning email and decommission. Document test results and any gaps.

Validation Gate

Check: Complete consumption architecture with catalog, governance, and end-to-end testing

Expected: T-shirt VM templates (S/M/L/XL) with Cloud Templates, service catalog published with categories, tenant onboarding checklist, resource quotas and lease policies configured, multi-level approval workflows, NSX VPC self-service integration, cost showback operational, Day-2 operations catalog, end-to-end test completed

Common Errors

Cloud Template deploys VM to wrong workload domain
Fix: Placement is controlled by Cloud Zones and constraints. Verify: (1) Cloud Zones are correctly mapped to workload domain clusters, (2) constraint tags in the Cloud Template match the Cloud Zone tags, (3) project is assigned to the correct Cloud Zones.
Approval workflow not triggering for production deployments
Fix: Approval policies require correct condition matching. Verify: (1) the condition references the correct input property (e.g., 'input.environment == production'), (2) the policy is assigned to the correct project, (3) the approver group has valid members with email notifications configured.
Resource quota exceeded but deployment still succeeds
Fix: Quota enforcement must be configured as 'Hard' limit (not 'Soft' warning). Check: Policies → Resource Quota → Enforcement Type. Also verify the quota scope matches the project scope.
Custom naming produces duplicate VM names
Fix: Ensure naming template includes a unique index: ${resource.count.index} or ${timestamp}. Also configure VCF Automation naming policy to reject duplicates. Common issue when deploying multiple VMs from the same catalog request.

Final Validation

Self-service consumption architecture operational with governance and cost transparency

✓ Service catalog → Published with VM, networking, Kubernetes, and Day-2 categories

✓ Cloud Templates → Parameterized templates deploying VMs with correct sizing, naming, and placement

✓ Tenant onboarding → Documented checklist covering project, quotas, roles, policies

✓ Governance → Resource quotas, lease policies, naming conventions, tag enforcement active

✓ Approval workflows → Multi-level approvals configured and tested

✓ Cost transparency → Per-project showback with resource rates and VCF Operations integration

Cleanup / Restore

• Delete test VMs and projects

• Revert approval policies to original state

• Save Cloud Template YAML for reuse

Design Reflection (VCDX)

Consumption architecture demonstrates your ability to transform infrastructure into a service. VCDX panelists will challenge: 'How do you prevent sprawl without killing developer velocity?' Answer with governance (quotas + leases + approvals) calibrated to organizational maturity. 'What happens when a project hits quota during a production incident?' Answer with emergency override process and post-incident quota review.

Requirements

  • Self-service VM provisioning in < 15 minutes for development teams
  • Resource governance preventing uncontrolled infrastructure growth
  • Cost transparency for per-team infrastructure spending

Constraints

  • VCF Automation is required (not optional) for self-service
  • Approval workflows must integrate with existing email/ticketing system
  • Naming conventions must comply with organizational CMDB standards

Assumptions

  • Development teams will adopt self-service catalog over ticket-based requests
  • T-shirt sizes cover 90%+ of workload profiles (custom sizing is exception, not rule)
  • Cost rates accurately reflect actual infrastructure costs

Risks

  • Low catalog adoption if templates don't match actual workload needs — mitigate with user feedback loop
  • Quota limits too restrictive → teams bypass catalog via direct vCenter access — mitigate with RBAC restricting direct provisioning
  • Lease policy deletes VMs with active workloads — mitigate with pre-expiry notifications and grace period

Self-Assessment Discussion Prompts

  1. How do you handle workloads that don't fit T-shirt sizes without creating sprawl of custom templates?
  2. When should approval be automated vs manual? What criteria determine the boundary?
  3. How do you measure catalog adoption and identify friction points in the self-service experience?
  4. What's the organizational change management needed to shift from ticket-based to self-service provisioning?

Extensions

Build a Terraform provider integration for teams that prefer IaC over catalog-based provisioning

Implement ABX extensibility actions for custom post-deployment configuration (join AD domain, install monitoring agent)

Design a multi-cloud catalog that provisions VMs on both VCF and AWS from the same template

Create a VPC-as-a-service catalog item that provisions isolated network environments on demand

⚠ Known Pitfalls (from Community KB)

Publishing templates without testing end-to-end — a broken catalog item destroys user trust in self-service
Setting quotas too tight initially — users abandon the catalog and go back to tickets; start generous and tighten based on actual usage
Ignoring Day-2 operations — provisioning is 10% of the lifecycle; resize, snapshot, and decommission workflows matter more for adoption
Not mapping cost rates to actual infrastructure costs — showback numbers that don't match finance reports lose credibility

References

  • VCF Automation Administration Guide: techdocs.broadcom.com
  • VCF Automation Cloud Template Reference: techdocs.broadcom.com
  • NSX VPC and Projects Guide: techdocs.broadcom.com
Was this page useful?
Type to search. ↑ ↓ to move, Enter to open, Esc to close.