Onboarding an Enterprise onto a New Stack

Onboarding an Enterprise onto a New Stack

Tech Stack

Terraform CloudGoogle CloudGitHub

Early on as a consultant at Insight, my team and I were tasked with onboarding an existing enterprise onto a number of platforms in a thoughtful and systematic way. The new stack included Terraform Cloud, Google Cloud Platform, Terraform as an IaC tool, and GitHub as the primary git version control system. None of these tools were being used at scale within the organization individually, and there wasn't widespread experience or exposure to any of them within the teams we supported, so the task was to learn these platforms, build a robust governance and observability framework that enabled and didn't block, figure out the best way for them to interact with each other, understand how to implement them at scale, and train and onboard teams across the stack.

Naturally, the intention of this stack was to enable teams to move from click-ops to a more robust, trackable solution for their infrastructure management. Github repos would be the source of truth for their infra code, Terraform Cloud would host their Terraform state file, and Terraform HCL would be their sole IaC solution.

Anyone whose worked on a project like this, though, understands all too well that a stack of expensive SaaS products doesn't a solution make. Our team was tasked with driving ease of use and adoption among app teams across the organization, and effort that took a remarkable amount of thought and solutioning. Turning a disjointed grouping of products into something that would accelerate and strengthen the infra deployment process across 300+ applications was a true learning experience.

Starting Principles

Firstly, we had to figure out how to govern each tool individually with regards to access, and also how each tool would reliably interact with others programmatically. Before diving into either, we set a few standards we knew our solutions needed to follow, including the standard that everything had to be git-backed, automated, repeatable, and programmatic. In practice, that meant when a user on an app team needed something, there was already a clear path they could follow and a clear pattern they fell into.

Finding the Lines So We Could Color Inside

Before we could design any governance model, we had to understand the organizational units of each platform and how configurations mapped to them. Teams, repositories, and orgs in GitHub; projects, teams, workspaces, and variable sets in Terraform Cloud; projects, folders, workload identity federation, service agents, and shared networking services in GCP.

We then had to assemble a catalog of the constraints each platform imposed. That process started with naming conventions: which characters were allowed (including what a name could start and end with), whether names were case sensitive and where, whether they needed to be globally unique or only unique within an org, and whether they could be changed after creation. GCP project IDs, for example, are limited to 6 to 30 characters, must be lowercase, must be globally unique, and are permanent once assigned. Terraform Cloud workspaces and GitHub repositories, on the other hand, allow longer names, only need to be unique within their org, and can be renamed, because each one is assigned an immutable internal ID at the time of creation. Workload Identity Federation added another wrinkle. The google.subject attribute has a hard limit of 127 bytes, and since Terraform Cloud builds the token's subject from a combination of the organization, project, workspace, and run phase, long workspace names would push it past that limit. That effectively capped how long Terraform Cloud workspace names could be.

Because a single identifier often travels across all three platforms (a GCP project backed by a Terraform Cloud workspace, whose code lives in a GitHub repository), the most restrictive rule on any one platform would effectively became a limitation for all of them. And where that wasn't practical, we had to document how the names related to each other.

Since Terraform would be our source of truth, we codified these naming limitations as regex patterns. I've shared some of them below.

We also documented structural limits, such as how many projects a single folder can hold, how deeply folders can be nested, and how many workspaces or teams an organization can support. Those ceilings would dictate how far any hierarchy we designed could scale before it'd have to be reorganized. One handy piece of information that isn't widely shared is that there's essentially no limit to how many workspaces a single Terraform Cloud org can host. At least, I haven't found the operational ceiling yet. The same can't be said for how many resources a single workspace's state file can manage. Once you get above 2,500 resources, things start to get cumbersome, and past 5,000, you're bound to see plans and applies degrade with wait times moving towards 10-15 minutes.

Finally, we mapped the metadata options available on each platform and where their limits sat. GCP offers both labels and tags, and they behave very differently. Labels are simple key-value pairs attached to individual resources and aren't inherited, while tags are centrally defined, inherited down the resource hierarchy, and can be referenced by IAM and organization policies. Terraform Cloud supports workspace tags for filtering and grouping, and GitHub offers repository topics and organization-level custom properties. Each metadata type came with its own caps on the number of keys, the length of keys and values, and the allowed character sets. Understanding these differences told us which platform could carry which piece of governance information, such as ownership, cost center, environment, or data classification, and where we'd need to enforce consistency ourselves because no single platform could do it for us. Again, where possible, we codified this into our Terraform. Those limits directly reshaped the design. We initially planned to nest folders under a single top-level folder per line of business in GCP, but the 300 folder-in-folder limit ruled that out. Rather than build a deeper LOB-based hierarchy to work around it, we went the other direction -- each app code got its own folder under a top-level folder, and we used round robin to distribute app folders across several top-level folders instead of grouping them by LOB.

Platform Limits at a Glance

GitHubGoogle Cloud PlatformTerraform Cloud
Org unit hierarchyOrganization > Team > Repository (flat)Organization > Folder (up to 10 levels deep) > ProjectOrganization > Project > Workspace
Naming charactersASCII letters, digits, . - _Project ID: lowercase letters, digits, hyphens; must start with a letter, no trailing hyphenWorkspace: letters, digits, hyphens, underscores
Naming lengthRepository: up to 100 charactersProject ID: 6 to 30 charactersWorkspace: 90 characters or fewer on its own, but when used with GCP Workload Identity Federation, the org, project, and workspace names combined can't exceed 78 bytes
Naming regex^[A-Za-z0-9._-]{1,100}$^[a-z][a-z0-9-]{4,28}[a-z0-9]$Workspace: ^[A-Za-z0-9_-]{1,90}$. Full WIF subject: ^(?=.{1,127}$)organization:[^:]+:project:[^:]+:workspace:[A-Za-z0-9_-]+:run_phase:(plan|apply)$
Case sensitivityCase is preserved in the name, but uniqueness and URLs are case-insensitiveLowercase onlyCase is preserved in the name
Uniqueness scopeUnique within the org or user accountProject ID: globally unique across all of GCPUnique within the org
RenamableYes, with automatic redirects from the old nameProject ID: no. Display name: yesYes
Immutable internal IDNumeric repository IDProject numberWorkspace ID (ws- prefix)
Structural limitsNone documented at this levelA folder can't hold more than 300 direct child folders; no per-folder project limit, but projects count against org and billing quotasNo practical workspace limit observed per org; state becomes cumbersome above about 2,500 resources per workspace
Labels / tagsTopics (up to 20 per repo, lowercase, up to 50 characters each) and org-level custom properties for key-value metadataLabels: up to 64 per resource, keys and values up to 63 characters, not inherited. Tags: centrally defined, inherited, usable in IAM and org policy conditionsTags up to 255 characters; letters, digits, colons, hyphens, underscores
Identity / federation constraintsOIDC token claims include repo and org namesgoogle.subject mapping capped at 127 bytes; WIF pool and provider IDs 4 to 32 charactersDynamic credentials build the token subject from org, project, workspace, and run phase, so long names can exceed GCP's 127-byte subject limit
Planning for Retries: Random Suffixes in Resource Names

For any resource type that goes into a pending-deletion or soft-delete state after it's removed, I tend to prefer to append a random ID to the end of the name. This will come up more often than you'd expect. GCP projects sit in a pending-deletion state for 30 days, and their IDs can't be reused, even after they've been fully purged. Azure subscriptions, Key Vaults, and several other resource types behave in similar ways. If part of your provisioning process fails partway through, often because proper idempotency hasn't been achieved just yet, a straightforward retry could get blocked by a resource still sitting in that to-be-deleted state under the name needed. Appending a random suffix means a failed or retried creation won't collide with whatever the last attempt left behind.

There's several other reasons to consider random suffixes, in moderation. Some resources, like GCP project IDs and Cloud Storage buckets, have to be globally unique, so a name that seems safe inside your org might already be taken elsewhere across GCP. The catch is that those characters aren't free. Every character in a random suffix cuts into your naming budget. You'll find that length budgets get tight quickly.

Personally, I've always been partial to random_pet. Instead of a hex string, it generates readable names like quick-otter, and they're much easier to spot in a console, a log, etc. Again, just keep an eye on length, since two or three words plus separators can eat into a tight limit fast.

Importing resources that use random IDs into Terraform used to be a little tricky, but there's a way around it. You can import a random_id resource using its b64_url value, with an optional prefix, which lets you bring an existing name under Terraform's management without triggering a diff.