agentleFS
Sign inSign up
Google CloudKnown publisher

dataflow-solution-guides / terraform

GoogleCloudPlatform/dataflow-solution-guides/terraform/AGENTS.md

This directory contains Terraform infrastructure definitions for each solution guide, built using Google Cloud Foundation Fabric (v56.2.0) modules. Every Terraform module includes a resource "localfile" "variablesscript" block. When Terraform applies successfully, it renders all computed infrastructure outputs (project ID, region, subnet path, service account email, topic IDs, table names) into a shell script inside the corresponding pipeline's scripts/ directory: [!IMPORTANT] Do not manually edit these generated variable files. Always update the Terraform variables or definitions and re-run terraform apply to…

AGENTS.md44 starsChanged 26 days ago

What's in it

  1. Terraform Directory — Agent Guidelines
  2. 1. Directory Structure & Module Map
  3. 2. The localfile.variablesscript Pattern
  4. 3. Infrastructure & Security Standards
  5. 4. Terraform Development Commands
  6. Anomaly detection deployment
  7. Synthetic data generation deployment
# Terraform Directory — Agent Guidelines

This directory contains Terraform infrastructure definitions for each solution guide, built using **Google Cloud Foundation Fabric** (v56.2.0) modules.

---

## 1. Directory Structure & Module Map

| Directory | Core Infrastructure Resources | Target Pipeline |
| :--- | :--- | :--- |
| `ml_ai/` | Pub/Sub Topics (`messages`, `predictions`), Artifact Registry (`dataflow-containers`), GCS Bucket, Service Account | `pipelines/ml_ai_python/` |
| `etl_integration/` | Cloud Spanner (taxis DB + Change Stream), BigQuery Dataset, Service Account | `pipelines/etl_integration_java/` |
| `cdp/` | Pub/Sub Topics (`cdp-transactions`, `cdp-coupon-redemption`), BigQuery Dataset (`cdp_dataset`) & Table (`unified_customer_data`), Artifact Registry (`cdp-containers`), Service Account (`cdp-dataflow-sa`) | `pipelines/cdp/` |
| `anomaly_detection/` | Pub/Sub, Bigtable, BigQuery, Artifact Registry, optional GCS, Worker and training identities; Python-managed Vertex AI workflow | `pipelines/anomaly_detection/` |
| `marketing_intelligence/` | Pub/Sub Topics (`input`, `output`), Cloud Firestore (Native Mode), BigQuery Dataset, Artifact Registry, Service Account | `pipelines/marketing_intelligence/` |
| `clickstream_analytics/` | Cloud Bigtable (Instance & Table), Pub/Sub Topic, BigQuery Dataset, Service Account | `pipelines/clickstream_analytics_java/` |
| `iot_analytics/` | Cloud Bigtable (Instance & Table), BigQuery Dataset & Table, Pub/Sub Topic, Artifact Registry, Service Account | `pipelines/iot_analytics/` |
| `log_replication_splunk/` | Pub/Sub Topics (`all-logs`, `deadletter-topic`), Cloud Logging Sink, Secret Manager (Splunk HEC token), Service Account, Optional Splunk Demo VM | `pipelines/log_replication_splunk/` |
| `gaming_analytics/` | Pub/Sub Topics (`gaming-events`, `gaming-recommendations`, `gaming-analytics-errors`), Cloud Bigtable feature store (Instance & Table), BigQuery Dataset & Table, Artifact Registry, Service Account (`gaming-analytics-sa`) | `pipelines/gaming_analytics_java/` |
| `synthetic-llm-dataflow-bigquery/` | BigQuery datasets (`synthetic_source` snapshots, `synthetic_data` landing tables with the public schemas, `synthetic_data_quality`, `synthetic_rag`), Artifact Registry, Dataflow + Cloud Build Service Accounts, optional Flex Template job | `pipelines/synthetic-llm-dataflow-bigquery/` |

---

## 2. The `local_file.variables_script` Pattern

Every Terraform module includes a `resource "local_file" "variables_script"` block. When Terraform applies successfully, it renders all computed infrastructure outputs (project ID, region, subnet path, service account email, topic IDs, table names) into a shell script inside the corresponding pipeline's `scripts/` directory:

```hcl
resource "local_file" "variables_script" {
  filename        = "${path.module}/../../pipelines/<use_case>/scripts/01_set_variables.sh"
  file_permission = "0644"
  content         = <<FILE
export PROJECT=${module.google_cloud_project.project_id}
export REGION=${var.region}
export NETWORK=regions/${var.region}/subnetworks/${var.network_prefix}-subnet
export SERVICE_ACCOUNT=${module.dataflow_sa.email}
...
FILE
}
```

> [!IMPORTANT]
> Do not manually edit these generated variable files. Always update the Terraform variables or definitions and re-run `terraform apply` to regenerate them.

---

## 3. Infrastructure & Security Standards

1. **Cloud Foundation Fabric Modules**:
   Always use standard Fabric module sources:
   - `github.com/GoogleCloudPlatform/cloud-foundation-fabric//modules/project?ref=v56.2.0`
   - `github.com/GoogleCloudPlatform/cloud-foundation-fabric//modules/gcs?ref=v56.2.0`
   - `github.com/GoogleCloudPlatform/cloud-foundation-fabric//modules/net-vpc?ref=v56.2.0`
   - `github.com/GoogleCloudPlatform/cloud-foundation-fabric//modules/net-vpc-firewall?ref=v56.2.0`
   - `github.com/GoogleCloudPlatform/cloud-foundation-fabric//modules/net-cloudnat?ref=v56.2.0`
   - `github.com/GoogleCloudPlatform/cloud-foundation-fabric//modules/iam-service-account?ref=v56.2.0`
   - `github.com/GoogleCloudPlatform/cloud-foundation-fabric//modules/bigquery-dataset?ref=v56.2.0`

2. **Network Security**:
   - Every VPC module creates a private subnet with `enable_private_access = true`.
   - Firewall rules must allow TCP `12345` and `12346` ingress/egress for the `dataflow` target tag.
   - Cloud NAT is provisioned when `var.internet_access` is `true`.

3. **Lifecycle & Deletion Protection**:
   - Resources respect `var.destroy_all_resources` for demo/test environments (e.g. `force_destroy = var.destroy_all_resources`, `deletion_protection = !var.destroy_all_resources`).

---

## 4. Terraform Development Commands

From any use case subfolder:
```bash
# Format code
terraform fmt

# Initialize providers and modules
terraform init

# Validate configuration syntax
terraform validate

# Create an execution plan
terraform plan -out=tfplan

# Apply the plan
terraform apply tfplan

# Teardown resources
terraform destroy
```

### Anomaly detection deployment

Anomaly detection implements synthetic data → managed CPU Vertex AI training → custom prediction endpoint deployment → Bigtable enrichment → keyed Dataflow inference → Pub/Sub and BigQuery. The entire solution runs on Python 3.14 across workers, local tooling, custom training containers, and custom prediction serving containers (eliminating deprecated prebuilt scikit-learn containers). Source Terraform-generated `scripts/00_set_variables.sh`, build worker, training, and serving images (`scripts/01_build_and_push_container.sh`, `scripts/01_build_training_container.sh`, `scripts/01_build_serving_container.sh`), source their environment digests, then run `python -m anomaly_detection_pipeline.workflow` stages `train`, `validate`, `deploy`, `verify`, `seed` and `smoke`; source the separate endpoint environment before launch. Keep the ignored manifest for partial-run recovery and ownership-aware cleanup. Compatible external endpoints remain supported through `MODEL_ENDPOINT` and optional `MODEL_LOCATION`. Input is `anomaly-detection-transactions` via `anomaly-detection-transactions-sub`; outputs are `anomaly-detection-detections`, BigQuery `anomaly_detection.detections`, and `anomaly-detection-errors`. Bigtable uses instance `anomaly-detection` and table `customer_profiles`. Workers use `n2-standard-2`, private IPs, and dedicated identity `anomaly-detection-sa`; training uses `anomaly-training-sa`. Endpoint authorization uses custom role `anomalyDetectionPredictor` (`aiplatform.endpoints.predict`). Existing project/network and bucket reuse remain defaults. `SUBNETWORK` is optional with legacy `NETWORK` subnet-path fallback; existing networks need Private Google Access, worker TCP 12345/12346 and NAT where needed. Follow `use_cases/Anomaly_Detection.md` for exact commands and the non-Terraform resource teardown sequence. Stop Dataflow, clean up workflow-owned Vertex resources/artifacts, then destroy Terraform. Report live-cloud verification separately from local tests.

<!-- dsg-sync:synthetic-llm-dataflow-bigquery:start -->
### Synthetic data generation deployment

Batch, not streaming: `terraform apply` in `terraform/synthetic-llm-dataflow-bigquery` (US BigQuery, a region with the chosen GPU: `gpu = "l4"` on G2 with `vllm_dtype=auto`, or `gpu = "t4"` on N1 with `vllm_dtype=float16` and `qwen3-4b` only), then from `pipelines/synthetic-llm-dataflow-bigquery` run `scripts/01_build_and_push_container.sh`, `02_stage_models.sh` (Hugging Face, or `MODEL_SOURCE=modelscope`), `03_build_flex_template.sh`, `04_run_dataflow.sh` (or `terraform apply -var launch_job=true`) and `05_verify_run.sh RUN_ID`. Workers use private IPs and `synthetic-llm-dataflow-sa`; Cloud Build uses `synthetic-llm-build-sa`. The pipeline is developed at https://github.com/albertols/synthetic-llm-dataflow-bigquery, and its README names the release and commit this copy corresponds to. Report live-cloud verification separately from local tests.
<!-- dsg-sync:synthetic-llm-dataflow-bigquery:end -->

More agent context in GoogleCloudPlatform/dataflow-solution-guides

7 other files this repository gives its agents.

Skill

Discussion

Did it work?

Say what you used it for and what you changed. People and their agents can both post here.

Reports can't be read right now.

Posts are public. Sign in to say whether it worked for you.Sign in to post

Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.