Self-Service Portal Design (Catalog + Projects)
Objectives
- Design a project hierarchy with organizational isolation and resource quotas
- Create catalog items with governance constraints (security, sizing, encryption)
- Design RBAC matrix with role definitions and permission boundaries
- Plan quota management and approval workflows for resource consumption
- Integrate self-service provisioning with NSX security policies and vSAN storage policies
Prerequisites
Access to VCF 9.0 and Aria Automation documentation
Prior labs: vcp-architect-01, vcp-architect-05, vcp-architect-06
Required skills:
- Aria Automation (formerly vRealize Automation) concepts
- RBAC and identity management
- Cloud template / blueprint design
- Resource quota management
Lab Environment
Design exercise — can deploy Aria Automation on Holodeck for hands-on validation
Tasks
Task 1 Project Hierarchy & Organizational Design
Design the project structure that maps organizational boundaries to VCF resource isolation with proper governance.
Design organization and project structure:
Aria Automation hierarchy:
- Organization (top level): Maps to business entity
- Projects: Resource and user boundary within the organization
- Cloud Zones: Infrastructure capacity pools mapped to projects
- Cloud Templates: Blueprints consumed by users within projects
Design for 600-VM environment with 3 business units:
Organization: CompanyName-VCF
├── Project: BU-Finance (CFO organization)
│ ├── Cloud Zone: prod-finance (WD-1 Tier-1 cluster, 20-VM quota)
│ ├── Cloud Zone: dev-finance (WD-2 Dev cluster, 30-VM quota)
│ └── Users: 15 developers, 3 admins, 1 project owner
├── Project: BU-Engineering (CTO organization)
│ ├── Cloud Zone: prod-engineering (WD-1 Tier-2 cluster, 150-VM quota)
│ ├── Cloud Zone: dev-engineering (WD-2 Dev cluster, 100-VM quota)
│ └── Users: 60 developers, 8 admins, 2 project owners
├── Project: BU-Operations (COO organization)
│ ├── Cloud Zone: prod-operations (WD-1 Tier-2 cluster, 50-VM quota)
│ └── Users: 25 operators, 5 admins, 1 project owner
└── Project: Platform-Team (Infrastructure)
├── Cloud Zone: mgmt (Management domain, restricted)
├── Cloud Zone: shared-services (cross-project infrastructure)
└── Users: 5 platform engineers, 2 adminsDesign resource quotas per project:
| Project | Max VMs | Max vCPU | Max RAM (GB) | Max Storage (TB) | Max Networks |
|---|---|---|---|---|---|
| BU-Finance | 50 | 200 | 800 | 10 | 5 |
| BU-Engineering | 250 | 1,000 | 4,000 | 50 | 15 |
| BU-Operations | 75 | 300 | 1,200 | 15 | 5 |
| Platform-Team | 30 | 120 | 480 | 20 | 10 |
| Total | 405 | 1,620 | 6,480 | 95 | 35 |
Over-subscription strategy:
- Total quotas (405 VMs) < capacity (600 VMs) = no contention at current allocation
- Allow soft quota overages of 10% with automatic alert to project owner
- Hard limits enforced by Aria Automation — provisioning fails if quota exceeded
- Quarterly review: Adjust quotas based on actual consumption (reclaim unused)
Quota enforcement:
- VM count: Hard limit (provisioning fails)
- vCPU/RAM: Hard limit with 10% soft buffer
- Storage: Warning at 80%, hard limit at 100%
- Networks: Hard limit (prevents segment sprawl)
Design cloud zone placement policies:
Cloud zones determine WHERE VMs are deployed:
- Cloud Zone: prod-finance
- Compute: WD-1 Tier-1 cluster (5 hosts)
- Storage policy: Tier-1-Critical (FTT=2, encryption enabled)
- Network: NSX segment finance-prod (172.16.30.0/24)
- Tags: placement:tier1, compliance:sox, department:finance
- Placement policy: Spread (distribute VMs across hosts for HA)
- Cloud Zone: prod-engineering
- Compute: WD-1 Tier-2 cluster (8 hosts)
- Storage policy: Tier-2-Business (FTT=1, IOPS limited)
- Network: NSX segment eng-prod (172.16.40.0/23)
- Tags: placement:tier2, department:engineering
- Placement policy: Binpack (consolidate for efficiency — dev workloads)
- Cloud Zone: dev-engineering
- Compute: WD-2 Dev cluster (4 hosts)
- Storage policy: Tier-3-DevTest (FTT=1, force provisioning)
- Network: NSX segment eng-dev (172.16.102.0/24)
- Tags: placement:dev, department:engineering
- Placement policy: Binpack (maximize utilization)
Capability tags:
- Catalog items specify required tags (e.g., 'compliance:sox')
- Cloud zone must have matching tag for deployment to succeed
- This prevents SOX-scoped VMs from landing on non-compliant infrastructure
Design project lifecycle management:
- Project onboarding:
- Request: New project via ITSM ticket (ServiceNow integration)
- Approval: Infrastructure team reviews capacity, creates project in Aria Automation
- Provisioning:
a. Create project with initial quota
b. Map cloud zones based on workload classification
c. Assign users and roles
d. Deploy initial network segments (Aria Automation → NSX API)
e. Provide catalog access
- SLA: 3 business days for new project creation
- Project resource review:
- Monthly: Automated utilization report per project
- Quarterly: Project owner reviews resource consumption vs quota
- Annual: Full project audit — decommission unused resources
- Project decommissioning:
- Trigger: Business unit request or resource audit finding
- Process:
a. Notify project users (30-day notice)
b. Export/migrate required VMs to other projects
c. Delete remaining VMs and networks
d. Release quota back to capacity pool
e. Archive project configuration for audit trail (90-day retention)
Validation Gate
Check: Project hierarchy designed with quotas, cloud zone mappings, and lifecycle procedures
Expected: 4+ projects with quota allocations, cloud zone placement policies with capability tags, project onboarding/review/decommissioning procedures documented
Common Errors
Task 2 Catalog Item Design with Governance Constraints
Create catalog items (cloud templates) that enforce security, sizing, and compliance policies while providing developer-friendly self-service.
Design catalog item taxonomy:
| Catalog Item | Purpose | Sizes Available | Security Features | Target Cloud Zone |
|---|---|---|---|---|
| Linux Web Server | NGINX/Apache web tier | S(2/4), M(4/8), L(8/16) | vTPM, DFW web-tier tag | Any production |
| Linux App Server | Java/Python app tier | S(2/8), M(4/16), L(8/32) | vTPM, DFW app-tier tag | Any production |
| Database Server | MySQL/PostgreSQL | M(4/16/200G), L(8/32/500G), XL(16/64/1T) | vTPM, encryption, DFW db-tier tag | Tier-1 only |
| Windows Server | General purpose | S(2/4), M(4/8), L(8/16) | vTPM, encryption at rest | Any |
| Dev Sandbox | Development VM | S(2/4), M(4/8) | DFW dev tag, no vTPM | Dev only |
| Kubernetes Cluster | TKG cluster | Small(3), Medium(5), Large(9) worker nodes | NSX networking, DFW k8s tag | Engineering |
Key sizes (vCPU/RAM in GB/Storage in GB where listed):
- S: 2 vCPU, 4-8 GB RAM
- M: 4 vCPU, 8-32 GB RAM
- L: 8 vCPU, 16-64 GB RAM
- XL: 16 vCPU, 64 GB RAM (requires approval)
Design a detailed cloud template (Linux App Server example):
Aria Automation Cloud Template (YAML-based):
name: Linux App Server
version: 2.1
inputs:
size:
type: string
enum: [Small, Medium, Large]
default: Medium
environment:
type: string
enum: [Production, Development]
application_name:
type: string
pattern: '^[a-z0-9-]{3,20}$'
owner_email:
type: string
lease_days:
type: integer
minimum: 1
maximum: 365
default: 90
resources:
app_server:
type: Cloud.vSphere.Machine
properties:
name: '${input.application_name}-app'
image: rhel9-hardened-golden
flavor: '${input.size}'
constraints:
- tag: 'placement:${input.environment == "Production" ? "tier2" : "dev"}'
networks:
- network: '${resource.app_network.id}'
storage:
bootDiskCapacityInGB: 80
security:
vtpm: true
secureboot: true
customizationSpec: linux-app-customization
tags:
- key: tier
value: app
- key: env
value: '${input.environment == "Production" ? "prod" : "dev"}'
- key: owner
value: '${input.owner_email}'
- key: lease_expiry
value: '${now() + duration(input.lease_days, "days")}'
app_network:
type: Cloud.NSX.Network
properties:
networkType: existing
constraints:
- tag: 'segment:app-${input.environment == "Production" ? "prod" : "dev"}'Governance constraints embedded:
- Image: Hardened golden image only (no custom ISOs)
- vTPM: Always enabled (non-negotiable for production)
- Tags: Automatically applied for DFW group membership
- Lease: Maximum 365 days; expiry triggers decommission workflow
- Naming: Regex-enforced naming convention
- Placement: Constraint tags control which cloud zone
Design governance policies:
- Lease management:
- All VMs have a lease (default: 90 days, max: 365 days)
- 14-day warning before lease expiry
- 7-day final warning
- At expiry: VM powered off (not deleted) — owner has 30 days to renew
- After 30 days: VM deleted (with 7-day trash recovery)
- Exception: Production Tier-1 VMs can have indefinite lease with annual review
- Approval policies:
- Self-service (no approval): Small/Medium VMs in dev cloud zones
- Manager approval: Large/XL VMs in any cloud zone
- Infrastructure team approval: Any VM in Tier-1 cloud zone
- Security team approval: VMs requesting external network access
- Approval via: Aria Automation approval workflow → integrates with ServiceNow
- Day-2 actions available to users:
- Power on/off/restart
- Resize (within allowed sizes — cannot exceed project quota)
- Add disk (up to storage quota)
- Snapshot (max 3 snapshots, max 72 hours — auto-delete after)
- Reconfigure (add NIC, change network — requires approval if crossing security zones)
- Delete (immediate, no approval needed — user's own resource)
- Prohibited actions (even for project admins):
- Change storage policy (must match catalog item governance)
- Disable vTPM or encryption
- Move VM to different project
- Connect to unmanaged network (bypass NSX)
Design monitoring and chargeback:
- Resource metering:
- Aria Operations tracks: vCPU hours, GB-RAM hours, GB-Storage used, IOPS consumed
- Metric granularity: Hourly samples, daily aggregation, monthly billing report
- Chargeback model:
| Resource | Unit Cost | Billing Basis |
|---|---|---|
| vCPU | $0.02/hour | Allocated (not consumed) |
| RAM | $0.005/GB/hour | Allocated |
| Storage (Tier-1) | $0.15/GB/month | Provisioned |
| Storage (Tier-2) | $0.08/GB/month | Provisioned |
| Storage (Dev) | $0.04/GB/month | Provisioned |
| Network segment | $50/segment/month | Fixed |
- Showback dashboard (per project):
- Current month cost
- Cost trend (month-over-month)
- Top 10 most expensive VMs
- Idle VM detection (CPU <5% for 7 consecutive days → flag for decommission)
- Orphaned resources (disks without VMs, unused networks)
- Cost optimization recommendations:
- Right-sizing: VMs consistently using <25% of allocated vCPU → suggest downsize - Powered-off VMs: VMs powered off for >30 days → suggest deletion - Dev/test scheduling: Auto power-off dev VMs at 8 PM, power-on at 7 AM → 50% cost savings
Validation Gate
Check: Catalog items designed with governance constraints, approval policies, and chargeback model
Expected: 6+ catalog items with size options and constraints, cloud template with embedded security controls, lease/approval/day-2 policies defined, chargeback model with resource metering
Common Errors
Task 3 RBAC Matrix & Identity Integration
Design the role-based access control model that governs who can do what within the self-service portal and underlying infrastructure.
Design role hierarchy:
| Role | Scope | Aria Automation | vSphere | NSX | SDDC Manager |
|---|---|---|---|---|---|
| Cloud Consumer | Project | Deploy from catalog, day-2 actions on own VMs | No access | No access | No access |
| Project Admin | Project | Full project mgmt, quota view, template import | Read-only (own VMs) | Read-only | No access |
| Cloud Admin | Organization | All projects, template design, cloud zone mgmt | Administrator (workload domains) | Read-only | Read-only |
| Infrastructure Admin | Global | Full Aria admin | Administrator (all) | Administrator | Administrator |
| Security Admin | Global | Read-only | Read-only (audit) | DFW policy admin | Read-only |
| Auditor | Global | Read-only (all projects) | Read-only | Read-only | Read-only |
Least-privilege principle:
- Cloud Consumers cannot access infrastructure directly (no vCenter, no NSX)
- Project Admins can manage their project but cannot affect other projects
- Cloud Admins can design templates but cannot bypass security policies
- Only Infrastructure Admins can modify SDDC Manager and management domain
Design identity source integration:
- Active Directory integration:
- vCenter SSO: AD/LDAPS identity source added
- Aria Automation: AD/LDAPS for user authentication
- NSX Manager: AD/LDAPS for admin access
- SDDC Manager: Local admin + AD for operator access
- AD group mapping:
| AD Group | Aria Role | vCenter Role | NSX Role |
|---|---|---|---|
| VCF-CloudConsumers | Cloud Consumer | N/A | N/A |
| VCF-ProjectAdmins | Project Admin | ReadOnly | N/A |
| VCF-CloudAdmins | Cloud Admin | Administrator | ReadOnly |
| VCF-InfraAdmins | Infra Admin | Administrator | Enterprise Admin |
| VCF-SecurityAdmins | ReadOnly | ReadOnly | Security Admin |
| VCF-Auditors | ReadOnly | ReadOnly | Auditor |
- MFA requirement:
- All admin roles (Cloud Admin, Infra Admin, Security Admin): MFA required
- Cloud Consumers: MFA optional (depends on organization policy)
- Implementation: AD FS with MFA provider (Azure MFA, Duo, RSA)
- Service accounts:
- Aria Automation → vCenter: Dedicated service account (vcf-aria-svc@domain.com) - Aria Automation → NSX: Dedicated service account (vcf-aria-nsx-svc@domain.com)
- Permissions: Minimum required for provisioning operations
- Password rotation: 90-day automated rotation via secrets manager
Design access control for sensitive operations:
- Emergency access (break-glass):
- Scenario: AD is down, normal authentication unavailable
- Procedure:
a. Use local admin accounts (vcfadmin@vsphere.local for vCenter)
b. Credentials stored in sealed envelope in physical safe
c. Every break-glass access triggers audit alert
d. Post-incident: Change all local admin passwords
- Privileged access management:
- Admin sessions: Maximum 8-hour session duration, auto-logout
- API access: Token-based with 1-hour expiry, scope-limited
- SSH to ESXi hosts: Disabled by default; enabled per-host via lockdown mode exception
- ESXi Shell: Timeout after 900 seconds of inactivity
- Audit trail:
- All authentication events logged to Aria Operations for Logs
- vCenter tasks and events: Retained for 365 days
- NSX audit log: All DFW changes logged with user identity
- SDDC Manager audit: All lifecycle operations logged
- Report: Monthly access review report for SOC 2 evidence
Create design decision D-012 — Self-Service Access Model:
Decision: Aria Automation with project-based isolation and catalog-driven provisioning
Alternatives:
A) Direct vCenter access for all users:
- Pro: Familiar interface, full flexibility
- Con: No governance, no quotas, no audit trail for provisioning decisions
- Rejected: Cannot enforce golden image usage, VM tagging, or lease management
B) ServiceNow only (ITSM ticket for all provisioning):
- Pro: Full change management for every VM
- Con: 3-5 day lead time for VM provisioning; developers bypass with shadow IT
- Rejected: Unacceptable velocity for development teams
C) Aria Automation (Catalog-based self-service):
- Pro: Developer velocity (minutes to deploy), governance embedded in templates, quota enforcement
- Con: Additional Aria license cost, template maintenance overhead
- SELECTED: Balances speed with control
Justification: Catalog-driven model reduces VM provisioning from 3-5 days to 5-10 minutes while maintaining governance controls (golden images, vTPM, tags, leases, quotas).
Validation Gate
Check: RBAC matrix designed with role hierarchy, AD integration, and access control policies
Expected: 6 roles with cross-platform permissions, AD group mapping, MFA requirements, break-glass procedure, audit trail design
Common Errors
Task 4 Quota Exhaustion & Alerting Design
Design the monitoring and alerting framework for quota exhaustion, idle resource detection, and capacity governance.
Design quota monitoring alerts:
| Alert | Condition | Severity | Action | Notification |
|---|---|---|---|---|
| Quota Warning | Project >70% of any quota | Warning | Email project owner | |
| Quota Critical | Project >90% of any quota | Critical | Email project owner + cloud admin | Email + Slack |
| Quota Exceeded | Provisioning blocked | Critical | Project owner must decommission or request increase | Email + Slack + ticket |
| Cluster Capacity Warning | Cluster >75% CPU or RAM | Warning | Infrastructure team reviews | PagerDuty |
| Cluster Capacity Critical | Cluster >85% CPU or RAM | Critical | Block new provisioning, procurement trigger | PagerDuty |
| Storage Warning | vSAN datastore >70% used | Warning | Infrastructure team reviews | |
| Storage Critical | vSAN datastore >80% used | Critical | Restrict storage-heavy catalog items | PagerDuty |
Escalation:
- Warning: Informational, no action required for 30 days
- Critical: Must acknowledge within 4 hours; action plan within 24 hours
- Exceeded: Immediate — provisioning is blocked
Design idle resource detection:
- Idle VM criteria:
- CPU utilization <5% average for 14 consecutive days
- Network traffic <1 Mbps for 14 consecutive days
- No console or SSH sessions for 14 days
- Exceptions: Monitoring agents, backup targets, license servers (tag-based exclusion)
- Idle resource workflow:
Day 1: VM identified as idle → automated tag applied (lifecycle:idle) Day 1: Email notification to VM owner: 'Your VM appears idle' Day 14: If still idle → second notification: 'VM will be powered off in 7 days' Day 21: VM automatically powered off (not deleted) Day 51: If still powered off → notification: 'VM will be deleted in 30 days' Day 81: VM deleted → storage reclaimed
- Orphaned resource detection:
- Disks without VMs (detached VMDKs)
- NSX segments with 0 connected VMs
- Snapshots older than 72 hours
- Report: Weekly orphaned resource report to project admins
Design quota increase workflow:
- Request:
- User: Submits quota increase request via Aria Automation custom form
- Fields: Project, resource type, requested increase, business justification
- Approval chain:
- Level 1: Project owner reviews business justification (auto-approve if <10% increase)
- Level 2: Cloud admin verifies infrastructure capacity
- Level 3: Finance approval if chargeback impact >$X/month
- Fulfillment:
- Approved: Cloud admin updates project quota in Aria Automation
- Denied: Feedback provided with alternative options (right-size existing VMs, decommission idle)
- SLA: 2 business days for standard requests, 4 hours for emergency
- Capacity-triggered procurement:
- When cluster utilization consistently >75% for 30 days → trigger procurement workflow
- Procurement lead time: 8-12 weeks for new hosts
- Bridge solution: Enable HCI Mesh from underutilized cluster if available
- Decision point: Add hosts to existing cluster vs create new workload domain
Design self-service portal operational summary:
- Architecture diagram:
[Users] → [Aria Automation Portal]
↓
[Catalog Items / Templates]
↓
[Approval Policies]──→ [ServiceNow Integration]
↓
[Cloud Zones + Placement]
↓
[vCenter API]──→ [VM Provisioning]
[NSX API]───→ [Network + DFW Tags]
[vSAN]─────→ [Storage Policy Applied]
↓
[Aria Operations]──→ [Monitoring + Quota Tracking]
↓
[Chargeback Reports]──→ [Finance Dashboard]- KPIs:
- Provisioning time: <10 minutes (target) from request to VM ready
- Quota compliance: 100% — no VM exists outside a project
- Idle VM ratio: <5% of total VMs (measured monthly)
- Catalog adoption: >90% of VMs deployed via catalog (not manual)
- Continuous improvement:
- Monthly: Review catalog item usage — retire unused items, add new ones
- Quarterly: Review quota allocations vs consumption — right-size quotas
- Annual: User satisfaction survey — gather feedback on catalog items and approval speed
Validation Gate
Check: Quota monitoring, idle detection, and governance workflows designed
Expected: 7+ alert rules with severity and action, idle VM workflow with timeline, quota increase approval chain, operational KPIs defined
Common Errors
Final Validation
Complete self-service portal design with project hierarchy, catalog governance, RBAC, and operational monitoring
✓ Project hierarchy with quotas → 4+ projects with resource quotas, cloud zone mappings, lifecycle procedures
✓ Catalog items with governance → 6+ catalog items with embedded security controls, approval policies, lease management
✓ RBAC matrix complete → 6 roles across Aria/vCenter/NSX/SDDC Manager, AD integration, break-glass procedure
✓ Operational monitoring designed → Quota alerts, idle detection workflow, chargeback model, KPIs
Cleanup / Restore
• Save all self-service design documentation and RBAC matrices
• If using Holodeck: Remove Aria Automation test projects and revert
Design Reflection (VCDX)
Self-service design reveals operational maturity. VCDX panelists look for: (1) Is governance embedded in the template, not bolted on? Golden images, vTPM, tags, leases should be non-negotiable within the catalog item. (2) Is the RBAC model least-privilege? Cloud consumers should never touch vCenter directly. (3) Is there a feedback loop? Quota alerts, idle detection, and chargeback close the loop between provisioning and governance. (4) What happens when things go wrong? Break-glass access, quota exhaustion handling, and orphaned resource cleanup are the operational reality checks.
Requirements
- VM provisioning in <10 minutes (developer velocity)
- All VMs must be tagged for DFW group membership
- Resource consumption tracked and chargeable to business units
- SOC 2 audit trail for all provisioning and access events
Constraints
- All VMs must use approved golden images (no custom ISOs)
- vTPM mandatory for all production VMs
- Project quotas cannot exceed cluster capacity
- Aria Automation license required (additional cost)
Assumptions
- Active Directory available and reliable for authentication
- ServiceNow available for approval workflow integration
- Business units accept chargeback model for resource consumption
- Golden images updated monthly with latest security patches
Risks
- Catalog adoption <90% — developers bypass portal and request manual VM creation
- Quota sprawl — projects accumulate quota over time without review
- Golden image vulnerability — compromised base image affects all new VMs
- Aria Automation downtime blocks all new provisioning
Self-Assessment Discussion Prompts
- What happens if Aria Automation is unavailable — how do you handle emergency VM provisioning?
- How would you extend this design to support Kubernetes namespace provisioning alongside VM provisioning?
- If a business unit disputes their chargeback report, how do you investigate and resolve?
- What is the operational cost of maintaining the self-service portal — is it justified for a 600-VM environment?
Extensions
GitOps-Driven Infrastructure with Aria Automation
Integrate Aria Automation templates with Git repositories. Developers submit pull requests with cloud template changes; CI/CD pipeline validates, tests in dev cloud zone, and promotes to production catalog. Document the Git-to-catalog pipeline, approval gates, and rollback procedure.
Multi-Cloud Self-Service with Aria Automation
Extend the catalog to provision VMs on AWS/Azure alongside VCF. Design cloud-agnostic templates, cross-cloud networking (HCX), and unified chargeback across private and public cloud. Document the cloud zone configuration for AWS/Azure and the placement policy logic.
Compliance-as-Code with Aria Automation ABX
Implement compliance checks as ABX (Action Based Extensibility) actions that run at provisioning time. Design actions that verify: VM is tagged, storage policy meets compliance, network segment is in approved zone, and vTPM is enabled. Document the ABX action code and failure handling.
⚠ Known Pitfalls (from Community KB)
References
- VCF Automation 8.x Cloud Template Reference
- VCF Automation RBAC and Identity Management Guide
- VMware VCF 9.0 — Private Cloud Consumption with Aria Automation
- VCF Operations — Chargeback and Showback documentation
- VCF Automation — Approval Policies and Governance