One GCP project. Three repositories. Six state files, and no seventh: the Cloudflare zone in front of everything stays console-managed. That is the Terraform architecture behind a personal site, a honeypot, an analytics platform, and a monitoring stack. None of these services share a deploy cycle. None of them should share a blast radius.
The default path with Terraform is a single directory, a single state file, and a single terraform apply that touches everything. That works until it does not. A bad apply to a Cloud SQL configuration should not delete your Pub/Sub topics. Rotating an API key should not require replanning your DNS records. Destroying a honeypot for maintenance should not risk your analytics database. Flat Terraform couples everything, and coupling is where infrastructure breaks.
Three specific problems drove this architecture. First, blast radius: a single state file means a single point of failure for every resource in the project. Second, secrets in state. Terraform's sensitive = true attribute hides values from terminal output but writes them to the state file in plaintext. Every third-party API key you manage through Terraform is one gsutil cat away from exposure. Third, IAM sprawl. The default GCP service account has Editor permissions on the entire project. A compromised container can do anything the project allows.
Layering solves each problem with a structural constraint. Independent state files limit blast radius to a single operational domain. A split secret strategy (CLI-populated containers for external credentials, Terraform-managed versions for system-generated passwords) keeps third-party keys out of state entirely. Per-service accounts with narrowly scoped role grants mean a compromised site container can publish to one Pub/Sub topic and nothing else.
The cross-layer wiring uses terraform_remote_state with a defaults block pattern that allows layers to deploy in any order after the foundation exists. Outputs from one layer become the read-only inputs to the next. The dependency direction is always downward: platform to services to site, never reversed. The Cloudflare layer sits outside the GCP dependency graph entirely; it only needs to know hostnames.
The practical result: each layer plans and applies in under 30 seconds. A secret rotation does not trigger a 15-minute Cloud SQL replanning. Tearing down the honeypot takes one terraform destroy in a single directory while the site keeps running. Every secret has a named accessor, so Cloud Audit Logs answer the question of who read what and when. I run six Terraform layers across two repos, with a third repo as the record for the console-managed Cloudflare zone, but the primitives are the same ones a team would use across fifty services and five projects. The boundaries scale. The architecture does not change.
This is what I run in production. It is not a reference design or a best-practices document. It is the outcome of hitting every failure mode flat Terraform creates and deciding each one was worth preventing.
Six layers across two repos, plus the edge
The infrastructure splits across three repositories by operational domain. One handles the personal site, honeypot, and supporting services. One handles observability: web analytics and dashboards. The third is the record for the Cloudflare network layer (DNS, WAF, Access, redirects) for the public zone, and it holds no Terraform: the zone is managed in the Cloudflare console and checked by a read-only posture monitor. The two GCP repos carry all six layers.
Updated 2026-09-27. The first version said the Cloudflare repo was a scaffold that would be implemented next, as a seventh state file. On 2026-09-19 I decided against that: the zone stays console-managed, and the repo became a committed baseline with a read-only check every six hours instead. The layer count and the diagrams below now reflect that; the GCP layers are unchanged.
site-infra/terraform/
platform/ -> Cloud SQL, Pub/Sub, BigQuery, Secret Manager, Artifact Registry
services/ -> Honeypot, alerting service
site/ -> Cloud Run (site), DNS records, IAM
observability-infra/terraform/
shared/ -> Project APIs, monitoring, cross-module IAM
analytics/ -> Cloud SQL (analytics), Cloud Run (analytics app)
dashboards/ -> Cloud Run (Grafana)
network-infra/ (no Terraform: console-managed zone, posture baseline only)
All six live layers write to the same GCS bucket with prefix isolation:
gs://my-project-tf-state/
platform/default.tfstate
services/default.tfstate
site/default.tfstate
observability-infra/shared/default.tfstate
observability-infra/analytics/default.tfstate
observability-infra/dashboards/default.tfstate
The network repo keeps no Terraform state, so it has no prefix here.
Each backend configuration is identical except for the prefix:
# site-infra/terraform/platform/backend.tf
terraform {
backend "gcs" {
bucket = "my-project-tf-state"
prefix = "platform"
}
}
Deploy order within site-infra: platform, then services, then site. Destroy order: reversed. Site first, services second, platform last. Observability layers deploy independently from each other and from the site repo. The network layer has no dependency on any GCP state file. It only needs hostnames.
The ordering constraint is strict within site-infra because each layer reads state from the one before it. If you destroy platform before site, the site layer loses its remote state references and terraform plan fails. The observability layers are independent entirely: different Cloud SQL instances, different service accounts, no cross-repo state reads.
Why each boundary exists
Platform vs. services. Platform creates the data pipeline (Pub/Sub, BigQuery), the container registry, and the secret containers. Services creates the honeypot and alerting. Destroying the honeypot for maintenance should never touch the data pipeline that other services also write to.
Services vs. site. The public site needs the honeypot's URL for reverse-proxy trap routes. But the site must deploy even if the honeypot does not exist yet. The defaults block on the remote state reference makes this possible (covered below).
Analytics vs. dashboards. Each has its own Cloud SQL instance, service account, and secrets. Redeploying the dashboard layer should not trigger a Cloud SQL plan diff in the analytics database.
GCP vs. Cloudflare. Compute and network are independent enough to deserve separate lifecycles and separate credentials. A Cloudflare WAF rule change has no business sharing state with a Cloud SQL configuration, so the zone sits outside the GCP state entirely: console-managed, with a read-only posture check. The Cloudflare repo only needs to know hostnames; it does not read GCP remote state.
Cross-layer wiring with terraform_remote_state
Outputs are the API of each layer. The platform layer exports everything downstream layers need: project ID, Pub/Sub topic name, secret IDs, BigQuery dataset ID, Artifact Registry path. Downstream layers read them through terraform_remote_state data sources.
# site-infra/terraform/platform/outputs.tf
output "project_id" {
description = "The GCP project ID. Used by all downstream layers."
value = google_project.main.project_id
}
output "pubsub_topic_name" {
description = "Pub/Sub topic for visitor events. Site and honeypot publish to this."
value = google_pubsub_topic.visitor_events.name
}
output "third_party_api_secret_id" {
description = "Secret Manager secret ID for the third-party API key."
value = google_secret_manager_secret.third_party_api_key.secret_id
}
The consuming layer references these values through a remote state data source:
# site-infra/terraform/services/remote_state.tf
data "terraform_remote_state" "platform" {
backend = "gcs"
config = {
bucket = "my-project-tf-state"
prefix = "platform"
}
}
# Usage in resource blocks:
# data.terraform_remote_state.platform.outputs.project_id
The defaults block for optional dependencies
The site layer depends on the services layer for the honeypot URL. But during initial deployment, the services layer might not exist yet. Without handling this, terraform plan fails because it cannot read a state file that does not exist.
The defaults block solves this:
# site-infra/terraform/site/remote_state.tf
data "terraform_remote_state" "services" {
backend = "gcs"
config = {
bucket = "my-project-tf-state"
prefix = "services"
}
defaults = {
honeypot_url = ""
alerting_url = ""
}
}
When the services state file is missing or the output is absent, Terraform falls back to the default value. The site container handles a blank honeypot URL gracefully: trap endpoints return a static response instead of proxying to the honeypot. This means the dependency direction is strict (platform to services to site) but not fragile. You can bring layers up in any order as long as the platform layer exists first.
Two secret management strategies
The two GCP repos use fundamentally different approaches to secrets, and the difference is deliberate.
Container-only pattern (site repo)
The site repo creates only the secret containers in Terraform. No google_secret_manager_secret_version resources for third-party API keys. Secret values are populated manually via CLI after terraform apply:
# site-infra/terraform/platform/secrets.tf
resource "google_secret_manager_secret" "third_party_api_key" {
secret_id = "third-party-api-key"
project = google_project.main.project_id
replication { auto {} }
}
# Populate after apply. The -n flag prevents a trailing newline.
echo -n "your-api-key" | gcloud secrets versions add third-party-api-key \
--project=my-project --data-file=-
The echo -n discipline matters. A trailing newline in a secret value silently breaks authentication against APIs that do strict string comparison. This has caused hours of debugging on more than one occasion.
The platform layer creates a handful of these containers, one per third-party service the rest of the stack talks to. The pattern is the same regardless of which API: container in Terraform, value via CLI, never both in state.
Hybrid case: a service account key for cross-cloud auth
One secret in the site repo sits between the two patterns: a service account key used by another system (a Supabase BigQuery foreign data wrapper) to read from BigQuery. The Secret Manager container is created in Terraform. The key itself is generated outside Terraform and populated via the same CLI pattern as third-party keys, then rotated on a schedule.
The reason it is not fully Terraform-managed: putting a service account key in secret_version would still write the JSON key into the state file. Same problem as a third-party API key. The CLI populate path keeps the key out of state.
Terraform-managed pattern (observability repo)
The observability repo uses random_password and google_secret_manager_secret_version to generate and store system credentials:
# observability-infra/terraform/analytics/cloud_sql.tf
resource "random_password" "db_password" {
length = 24
special = false
}
resource "google_secret_manager_secret_version" "analytics_db_password" {
secret = google_secret_manager_secret.analytics_db_password.id
secret_data = random_password.db_password.result
}
This password exists nowhere outside the system. Terraform generated it, stored it in Secret Manager, and uses it to configure the Cloud SQL user. The value in state is a random string with no external significance. The same pattern applies to the analytics app secret and the dashboard layer's admin password.
The decision rule
If a secret has value outside your infrastructure (an API key issued by a third party, a service account key, a credential you could reuse elsewhere), populate it via CLI so it never touches Terraform state. If Terraform generates the secret and it only lives within your system (a random database password, an app signing key), Terraform can manage the version. The sensitive = true attribute on Terraform variables only hides values from terminal output. It does not prevent them from being written to the state file in plaintext. [^1]
IAM at the narrowest scope
Every service account gets exactly the permissions it needs, scoped to the narrowest resource possible.
The site's service account has exactly one role:
# site-infra/terraform/site/iam.tf
resource "google_service_account" "site" {
account_id = "site-runner"
display_name = "Site Cloud Run Service Account"
project = data.terraform_remote_state.platform.outputs.project_id
}
# Topic-scoped, not project-scoped. A project-level roles/pubsub.publisher
# would let a compromised site publish to every topic in the project.
resource "google_pubsub_topic_iam_member" "site_visitor_events_publisher" {
topic = data.terraform_remote_state.platform.outputs.pubsub_topic_name
role = "roles/pubsub.publisher"
member = "serviceAccount:${google_service_account.site.email}"
}
That is the shape of the entire IAM surface for the site container: no project-level roles at all. Every grant names the resource it applies to. Today that is publisher on two topics (visitor events and the recon-chain trigger), secretAccessor on the two honeypot salt containers, and run.invoker on the honeypot service it proxies to. If the site is compromised, the attacker can publish to those two topics, read those two salts, and call the honeypot. They cannot deploy containers, modify IAM policies, read other secrets, or touch BigQuery or Cloud SQL, because no grant exists, not because a grant is narrow.
Corrected 2026-09-23. The first version of this section showed a project-level publisher grant and claimed the site read no secrets. An external review in September 2026 caught the project-level grant, and the salts moved into Secret Manager the same month; the text now matches the deployed layer.
The honeypot still reads its own credentials with its own service account. Pushing secret access down to the service that needs it keeps the front-most container's blast radius small and auditable.
The observability repo applies the same principle at finer grain. Instead of project-level grants, each secret gets its own IAM binding:
# observability-infra/terraform/analytics/iam.tf
resource "google_secret_manager_secret_iam_member" "analytics_db_password" {
project = var.project_id
secret_id = google_secret_manager_secret.analytics_db_password.secret_id
role = "roles/secretmanager.secretAccessor"
member = "serviceAccount:${google_service_account.analytics.email}"
}
The analytics service account can read the analytics database password, the analytics app secret, and the analytics database URL. It cannot read the dashboard layer's admin credentials. Same principle as the site SA, applied at finer grain: scope access to the resource, not the project.
Split CI service accounts: planner vs. deployer
The observability repo's shared/ layer manages two CI service accounts that run Terraform itself, with deliberately different permissions.
# observability-infra/terraform/shared/iam.tf (paraphrased)
# github-planner: plan-only. Used by `terraform plan` on pull requests.
# Eight project-level viewer/reader roles, plus one bucket-level grant below.
resource "google_project_iam_member" "github_planner" {
for_each = toset([
"roles/bigquery.metadataViewer",
"roles/cloudsql.viewer",
"roles/iam.securityReviewer",
"roles/logging.viewer",
"roles/monitoring.viewer",
"roles/run.viewer",
"roles/secretmanager.viewer", # metadata only, not values
"roles/serviceusage.serviceUsageConsumer",
])
project = var.project_id
role = each.value
member = "serviceAccount:${var.planner_sa_email}"
}
# The planner's one write capability: objects in the state bucket, and only
# under this repository's state prefixes (locking and refresh need it).
resource "google_storage_bucket_iam_member" "tf_state_planner_objects" {
bucket = var.state_bucket
role = "roles/storage.objectUser"
member = "serviceAccount:${var.planner_sa_email}"
condition {
title = "this-repo-prefixes-only"
expression = "resource.name.startsWith(\"projects/_/buckets/${var.state_bucket}/objects/site-infra/\")"
}
}
# github-deployer: write. Used by `terraform apply` on merges to main.
# Granted ~14 admin roles, all required for the modules in this repo.
The threat model is explicit. A malicious or compromised pull request could include Terraform code that mutates infrastructure during plan-time refresh (some providers issue write-adjacent calls during refresh). If the plan job runs as the deployer SA, that mutation succeeds. If it runs as the planner, the mutation fails at the GCP authorization layer. The planner's one write capability is the state bucket, scoped by an IAM condition to the repository's own state prefixes, which locking and refresh require; a hostile plan can therefore corrupt versioned state, not infrastructure. The same planner identity serves both repositories, each with its own prefix condition, so the boundary is per prefix, not per identity. The split protects the plan step, not the merge decision: apply runs as the deployer, and a malicious PR that gets merged does ship. Review is still the gate.
This is one of the strongest patterns in the stack and one of the easiest to overlook. Most CI Terraform setups use a single service account for both plan and apply. Splitting them costs little and closes a real attack path.
Both SAs are managed by Workload Identity Federation. Neither has a long-lived JSON key. The federation binding ties the SA to a specific GitHub repository and branch.
Deletion protection and drift prevention
Three patterns prevent the most common Terraform accidents:
Deletion protection on stateful resources. Cloud SQL instances and BigQuery tables get deletion_protection = true. You must explicitly set this to false before you can destroy them. This catches the most expensive class of Terraform mistake: accidentally destroying a database during a refactor. The site repo's BigQuery tables and the observability repo's Cloud SQL instances all carry this guard.
resource "google_sql_database_instance" "analytics" {
# ...
deletion_protection = true
}
Provider version pinning with pessimistic constraints. The ~> operator allows patch updates but blocks major version changes that might include breaking API changes. The two GCP repos sit on different majors today: the site repo on ~> 5.0, the observability repo on ~> 7.28. They were initialized at different times and have different upgrade pressures. Pinning per-repo is intentional. A provider major upgrade in one repo does not force a coordinated upgrade across all six layers.
required_providers {
google = {
source = "hashicorp/google"
version = "~> 5.0" # site-infra
# version = "~> 7.28" # observability-infra
}
}
The .terraform.lock.hcl file records exact provider versions and checksums per layer. It is committed to version control so every apply (local or CI) uses the identical provider binary.
API enablement protection. Every google_project_service resource includes disable_on_destroy = false. Without this, removing an API enablement resource from your Terraform config disables the API project-wide. Any service that depends on it breaks, Terraform-managed or not.
resource "google_project_service" "cloud_run" {
service = "run.googleapis.com"
disable_on_destroy = false
}
What this enables
Independent deploys. Each layer plans and applies in under 30 seconds. A secret rotation in the platform layer does not trigger a 15-minute Cloud SQL plan in the services layer.
Safe teardowns. Destroying the honeypot for maintenance takes one terraform destroy in the services directory. The site keeps running. The data pipeline keeps collecting.
Auditable secret access. Every secret has a named accessor. Cloud Audit Logs record which service account read which secret and when. When something goes wrong, the investigation starts with a query, not a guess.
Bounded compromise. The site SA has one role. The analytics SA can read its own secrets and nothing else. The CI planner SA cannot mutate. None of these are exotic patterns. They just require treating "what is this principal allowed to do" as a question to answer per resource, not per project.
Portable patterns. I run six Terraform layers across two repos for one personal stack. A team running 50 services across five projects would use the same primitives. Prefix-isolated state, outputs as APIs, scoped IAM, split CI SAs, and the container-only vs. Terraform-managed secret split. The boundaries scale. The architecture does not change.
References
[^1]: Terraform sensitive variables documentation. The sensitive = true attribute suppresses values in CLI output but does not encrypt them in state files.