The GCP documentation presents Cloud Run + Cloud SQL as a three-step exercise: add --add-cloudsql-instances, set your socket path, deploy. That description holds until you touch IAM auth, non-trivial passwords, schema migrations, or controlled traffic routing. At that point, the real operational behavior surfaces - and the error messages almost never point to the actual cause.
I run this stack at home on production-grade workloads. What follows is an accounting of five failure categories I hit, what each error said, what actually broke, and what fixed it.
Error Messages Lie Here More Than Anywhere Else
The thread connecting all five problems is misdirection. A password parsing failure reports connection refused. A MySQL authentication incompatibility reports Access denied as if the password is wrong. A cold-start race condition reports a timeout as if the migration is slow. None of these error messages name the actual problem. That gap is what makes debugging this stack expensive.
The five failure categories, and their tested workarounds:
Passwords embedded in DATABASE_URL strings break silently. Any special character with syntactic meaning in a URI - @, #, $, % - splits or corrupts the parsed host, port, or credentials. The reported error is connection refused or host not found, which points at network, not string parsing. The fix: pass credentials as discrete parameters to your database driver. Never embed a password in a URL string. This also eliminates a class of framework-specific double-encoding bugs.
MySQL 8.4 broke unix socket authentication until proxy v2.19.0. MySQL 8.4 changed its handshake so the server always advertises caching_sha2_password, regardless of configured defaults. Older Cloud SQL Auth Proxy releases did not support this plugin over unix sockets. The failure reports ERROR 1045 (28000): Access denied - identical to a wrong password. This was tracked as a P0 in the proxy's GitHub issues (#2317) and fixed in v2.19.0. If you pin the proxy sidecar, pin v2.19.0 or later. If you cannot control the proxy version, use TCP or a Language Connector.
On-boot migrations fail under Cloud Run's scaling model. Running migrations inside the startup sequence creates two concurrent problems. Multiple instances race each other against the same schema version table, and Cloud Run's four-minute startup window kills a slow migration mid-execution. The correct pattern is Cloud Run Jobs - a blocking pipeline step that runs before the new revision deploys.
Cloud Run promotes new revisions immediately by default. If a migration adds a NOT NULL column, the old revision starts failing inserts the moment the migration runs. Cloud Run has no built-in primitive for "hold until migration succeeds, then promote." That orchestration requires --no-traffic with revision tags and an explicit update-traffic call in your pipeline.
Connection pools multiply with instances. Cloud Run scales horizontally, and each instance initializes its own pool. A pool size of 10 across 20 instances creates 200 active connections - often exceeding Cloud SQL's max_connections limit. The Auth Proxy does not pool; it passes each application connection through to the database. Fixing this requires either Cloud SQL Managed Connection Pooling, a PgBouncer sidecar, or aggressively small per-instance pool sizes.
What to Change
These five problems share a common pattern: the default configuration optimizes for simplicity of initial setup, not for correctness under load or non-trivial operational conditions. In this environment, the changes that eliminated most of the friction were:
- Pass credentials as discrete parameters; use Secret Manager with
--set-secrets - Run migrations as Cloud Run Jobs with
--waitin the pipeline, gated before revision deploy - Use
--no-trafficplus revision tags for any migration that changes column constraints - Set per-instance pool size to 1-5 in this deployment; use Cloud SQL Managed Connection Pooling for PostgreSQL
- Use the sidecar pattern with
container-dependenciesordering, or switch to Language Connectors for MySQL 8.4
The technical layer below covers each problem in detail: the error, the mechanism, the workaround, and the remaining open issues.
The Password Parsing Trap
URI parsers assign syntactic meaning to several characters. When a database password contains @, #, $, %, /, or +, a DATABASE_URL like postgresql://user:p@ss#[email protected]:5432/mydb misdirects the parser. The first @ splits userinfo from host. The parser reads p as the password and ss#[email protected] as the host. The resulting error - connection refused or host not found - points at network connectivity, not string parsing. [1]
Percent-encoding resolves the immediate problem: @ becomes %40, # becomes %23, $ becomes %24. However, some frameworks re-interpret percent-encoded strings. Symfony treats %40 as a Symfony parameter reference and throws The parameter "40" must be defined, requiring double-encoding or bypassing URL-based config entirely. [2]
The workaround that holds across all frameworks: pass credentials as discrete driver parameters.
# Avoid: password embedded in URL
DATABASE_URL = "postgresql://user:p@ss#[email protected]:5432/mydb"
# Prefer: discrete parameters
import psycopg2
conn = psycopg2.connect(
host="127.0.0.1",
port=5432,
dbname="mydb",
user="user",
password="p@ss#1"
)
Most PostgreSQL and MySQL client libraries support discrete parameters directly. This eliminates the parsing layer and the class of framework-specific encoding bugs that accompany it.
IAM Authentication and the Missing Password Field
Cloud SQL IAM database authentication works through the Auth Proxy's --auto-iam-authn flag. The Proxy handles token exchange automatically. Frameworks that require a non-empty password field will fail despite correct IAM configuration. The JDBC driver, for example, passes a blank password string and receives The server requested password-based authentication, but no password was provided by plugin null.
The fix: configure the driver to use the Cloud SQL Socket Factory or Language Connector directly rather than routing credentials through the Proxy's password field. For applications where the framework enforces non-empty passwords, a placeholder value in the password field combined with --auto-iam-authn resolves the mismatch without affecting IAM token exchange. [3]
The MySQL 8.4 Unix Socket Incompatibility
This problem carried a priority: p0 label in the Auth Proxy issue tracker until the v2.19.0 fix landed. [4] MySQL 8.4 changed handshake behavior so the server always advertises caching_sha2_password in the initial authentication exchange. Older Auth Proxy releases did not support this plugin over unix sockets. The failure looks like this:
ERROR 1045 (28000): Access denied for user 'my-user'@'cloudsqlproxy~34.124.164.49' (using password: YES)
Nothing in that message identifies unix sockets or caching_sha2_password. It looks like a wrong password.
The mechanism: MySQL 8.4 sends caching_sha2_password in the initial handshake regardless of configured defaults. In proxy releases before v2.19.0, the unix socket path did not support the RSA key exchange this plugin requires. The TCP path worked because the server cached authentication after a successful connection. When that cache expired after maintenance or a connection reset, the unix socket broke again silently. [5]
The MySQL bug tracker records this as a server-side behavior change in 8.4 (Bug #118447). The server sends caching_sha2_password in the initial handshake regardless of the configured default plugin.
Current Fixes and Workarounds
As of June 2026, the following approaches resolve this problem in this environment:
- Pin the Auth Proxy sidecar to v2.19.0 or later. That release added
caching_sha2_passwordsupport for proxy clients using unix sockets. - Switch to TCP by using the
--portflag instead of--unix-socket. TCP connections bypass the older unix socket authentication path. - Use Language Connectors (
cloud-sql-python-connector,cloud-sql-node-connector,cloud-sql-go-connector). These handled MySQL 8.4 authentication before the binary proxy fix and remain a viable path. - Stay on MySQL 8.0 or set
mysql_native_passwordas the default plugin only as a temporary deferral strategy.
The binary proxy fix merged in October 2025 and shipped with v2.19.0. Do not upgrade a Cloud Run + Cloud SQL MySQL deployment to 8.4 on unix sockets without first verifying the proxy version in the path you actually run.
Migration Execution Has No Safe Default Path on Cloud Run
Running schema migrations inside the container startup sequence - the default in most framework quickstart guides - creates compounding failures on Cloud Run.
Cloud Run scales instances horizontally on traffic spikes. Multiple instances can start simultaneously, each executing the same migration. Flyway and Liquibase use advisory locks to serialize execution, but Cloud Run's four-minute startup window kills a slow migration before it completes. The container fails its health check, Cloud Run terminates it, and the migration may be partially applied.
The three specific failure modes:
- Parallel cold starts race each other against the same schema version table, causing lock contention or duplicate execution
- Even a fast checksum scan adds latency to every cold start, degrading first-request response times across all instances
- A mid-migration container kill leaves the schema partially applied - incompatible with both the old and new revision
Cloud Run Jobs as the Migration Executor
Cloud Run Jobs run to completion, connect to Cloud SQL via the same --add-cloudsql-instances mechanism as services, and are not time-bounded by an HTTP health check. The pipeline sequence that works in this environment:
# Step 1: Build the migration container image
docker build -t gcr.io/my-project/my-migrations:$BUILD_ID .
docker push gcr.io/my-project/my-migrations:$BUILD_ID
# Step 2: Update the Cloud Run Job with the new image
gcloud run jobs update run-migrations \
--image=gcr.io/my-project/my-migrations:$BUILD_ID \
--region=us-central1
# Step 3: Execute the migration (blocking - pipeline waits)
gcloud run jobs execute run-migrations \
--region=us-central1 \
--wait
# Step 4: Deploy the application revision (only runs if Step 3 exits 0)
gcloud run deploy my-service \
--image=gcr.io/my-project/my-app:$BUILD_ID \
--region=us-central1
The --wait flag makes the pipeline step block until the job completes. This provides a clean gate: if the migration fails, the pipeline stops and the new revision never deploys. [6]
Using a separate migration-only image keeps the job focused and produces a clear audit log of what ran and when. The alternative - running manage.py migrate or db:migrate inside the full application image - works but carries unnecessary dependency weight.
Traffic Routing After Schema Changes Requires Manual Orchestration
Cloud Run's default deployment behavior promotes a new revision to 100% of traffic immediately. This breaks when a migration adds a NOT NULL column with no default. The old revision starts failing inserts the moment the migration completes, because the column exists and the old code does not populate it.
The --no-traffic and Revision Tag Pattern
The operational workaround uses --no-traffic to deploy the new revision without routing traffic to it, then promotes after the migration succeeds and the new revision passes a health check:
# Deploy new revision to receive zero traffic, assign a tag
gcloud run deploy my-service \
--image=gcr.io/my-project/my-app:v2 \
--no-traffic \
--tag=green \
--region=us-central1
# Execute migration (blocking)
gcloud run jobs execute run-migrations \
--region=us-central1 \
--wait
# Verify the tagged revision at its dedicated URL
curl https://green---my-service-<hash>-uc.a.run.app/health
# Cut all traffic to the new revision
gcloud run services update-traffic my-service \
--to-tags=green=100 \
--region=us-central1
Cloud Run has no built-in primitive that holds a revision until a job succeeds and then promotes automatically. That orchestration lives in Cloud Build or your CI system. One additional edge case matters: if a canary split is already active, deploying without --no-traffic inherits the existing split. The new revision may receive 0% traffic regardless. The deploy output does not make this behavior obvious. [7]
The Expand-Contract Pattern for Backward-Compatible Schema Changes
The root solution to migration and traffic coupling is maintaining backward-compatible schema changes across release boundaries:
- Never drop a column in the same release that stops writing it. Deprecate in release N, stop writing in N+1, drop in N+2.
- Add
NOT NULLcolumns with a default first, backfill the data, then remove the default in a follow-on release. - Use expand-contract: add the new column, dual-write in the application, migrate existing data, switch reads to the new column, remove the old column in a future release.
This approach lets migrations run against the live database without requiring simultaneous traffic stops, enabling the Cloud Run Jobs pattern without coordinated migration and revision promotion.
Connection Exhaustion Under Serverless Scaling
Cloud Run scales instances horizontally. Each instance initializes its own connection pool. An application configured with a pool size of 10 running across 20 instances creates 200 active connections. Cloud SQL's default max_connections for a small instance is often 100 or lower. The result under load:
FATAL: remaining connection slots are reserved for non-replication superuser connections
This error appears under load, not during development. The Auth Proxy does not act as a connection pooler. It handles TLS termination and authentication, but each connection from the application maps to one connection in Cloud SQL. [8]
The viable mitigations, in order of preference in this environment:
- Cloud SQL Managed Connection Pooling (PostgreSQL, GA in September 2025): built-in PgBouncer at the Cloud SQL layer, scales automatically based on instance vCPU count, no separate infrastructure to operate. [9]
- PgBouncer as a Cloud Run sidecar: pools at the session or transaction level, reducing Cloud SQL connections regardless of instance count.
- Per-instance pool size of 1-5: for Cloud Run, a pool size of 1-2 per instance is often sufficient given the instance-level concurrency model. Applications sized for long-lived VMs default to pools far larger than serverless instances need.
--max-instancescap: sets an upper bound on instance count and therefore on connection count, at the cost of potential latency under high concurrency.
Sidecar vs. Built-in Socket Injection vs. Language Connector
Cloud Run supports multi-container deployments, enabling the Auth Proxy as a sidecar rather than using the built-in --add-cloudsql-instances socket injection. Table 1 compares the three approaches.
Table 1 compares Cloud SQL Auth Proxy connectivity options for Cloud Run.
| Approach | Version Control | Log Visibility | MySQL 8.4 Safe | Complexity |
|---|---|---|---|---|
--add-cloudsql-instances (built-in) |
None - GCP-managed | Proxy logs merged with service | Verify platform version | Lowest |
| Auth Proxy sidecar | Pin to any version | Separate container logs | Yes, with v2.19.0+ | Medium |
| Language Connector (in-process) | Per-language package | Application logs | Yes | Medium |
The sidecar pattern gives full control over the proxy version, flags (--auto-iam-authn, --private-ip, --max-signers), and startup ordering. The run.googleapis.com/container-dependencies annotation ensures the Proxy container passes its TCP health check before the application container starts. This eliminates the race condition where the application attempts a database connection before the Proxy is ready.
# Excerpt: Cloud Run service YAML with sidecar and startup ordering
apiVersion: serving.knative.dev/v1
kind: Service
metadata:
name: my-service
annotations:
run.googleapis.com/launch-stage: BETA
spec:
template:
metadata:
annotations:
run.googleapis.com/container-dependencies: '{"app":["cloud-sql-proxy"]}'
spec:
containers:
- name: app
image: gcr.io/my-project/my-app:latest
env:
- name: DB_HOST
value: "127.0.0.1"
- name: DB_PORT
value: "5432"
- name: cloud-sql-proxy
image: gcr.io/cloud-sql-connectors/cloud-sql-proxy:2
args:
- "--port=5432"
- "my-project:us-central1:my-instance"
startupProbe:
tcpSocket:
port: 5432
initialDelaySeconds: 2
periodSeconds: 5
The Language Connector approach (cloud-sql-python-connector, cloud-sql-node-connector, etc.) eliminates the proxy binary entirely. Connectivity runs in-process using the same mTLS and OAuth2 mechanism as the proxy. This remains a good path for MySQL 8.4 and for teams that prefer to avoid operating a proxy sidecar.
IAM Authentication Edge Cases
Enabling --auto-iam-authn introduces three operational edge cases not covered in the primary documentation.
Certificate refresh race conditions caused silent hangs in earlier versions. A background refresh could overwrite a valid certificate, leaving the Proxy stuck. Connections failed with Cloud SQL IAM service account authentication failed and the Proxy required a restart. Application-level retry logic did not help. This was fixed in proxy v2.11.3 (May 2024), but the pattern illustrates why monitoring proxy health matters. [10]
The IAM DB username derives from the service account email. For service accounts, the database username is the SA email minus the .gserviceaccount.com suffix, truncated at 63 characters for PostgreSQL. A mismatch here produces Access denied - identical to a wrong password, with no indication the username derivation is the issue.
GOOGLE_APPLICATION_CREDENTIALS overrides gcloud auth. When this environment variable is set, the Proxy uses the key file rather than the active gcloud identity. Running locally against a dev Cloud SQL instance with this variable set for another service silently authenticates as the wrong identity. Clear the variable or set it explicitly to the correct key before local testing.
What Actually Works: A Summary
The six changes that eliminated most of the operational friction in this environment:
- Credentials: Pass as discrete driver parameters. Use Secret Manager with
--set-secretsfor injection at runtime. - Migrations: Run as Cloud Run Jobs with
--waitin the CI pipeline. Gate the revision deploy on job exit code. - Traffic control: Deploy with
--no-trafficand a revision tag. Promote with an explicitupdate-trafficcall after validation. - Connections: Use Cloud SQL Managed Connection Pooling (PostgreSQL) or a PgBouncer sidecar. Set per-instance pool size to 1-5 in this deployment.
- Proxy setup: Use the sidecar pattern with
container-dependenciesstartup ordering, or switch to a Language Connector for MySQL 8.4. - Schema changes: Follow the expand-contract pattern. Maintain backward compatibility across revision boundaries.
References
[1] Stack Overflow. "How to handle special characters in a PostgreSQL URL connection string." https://stackoverflow.com/questions/23353623/how-to-handle-special-characters-in-the-password-of-a-postgresql-url-connection
[2] Symfony GitHub issue #26721. "Percent-encoded special characters for database password." https://github.com/symfony/symfony/issues/26721
[3] Google Cloud. "Overview of Cloud SQL IAM database authentication." https://cloud.google.com/sql/docs/postgres/iam-authentication
[4] GoogleCloudPlatform/cloud-sql-proxy. "Support caching_sha2_password authentication plugin with Proxy in Unix Socket mode." Issue #2317. https://github.com/GoogleCloudPlatform/cloud-sql-proxy/issues/2317. See also: "Cloud SQL Auth Proxy v2.19.0 release notes." https://github.com/GoogleCloudPlatform/cloud-sql-proxy/releases/tag/v2.19.0
[5] MySQL Bug #118447. "MySQL 8.4 always uses caching_sha2_password in the initial handshake." https://bugs.mysql.com/bug.php?id=118447
[6] Google Cloud Blog. "Running database migrations with Cloud Run Jobs." https://cloud.google.com/blog/topics/developers-practitioners/running-database-migrations-cloud-run-jobs/
[7] Google Cloud. "Rollbacks, rollouts, and traffic migration on Cloud Run." https://cloud.google.com/run/docs/rollouts-rollbacks-traffic-migration
[8] SADA Engineering Blog. "Cloud SQL Auth Proxy demystified." https://engineering.sada.com/cloud-sql-auth-proxy-demystified-54dd803d1474
[9] Google Developer Community. "Optimizing performance and scaling with Managed Connection Pooling for Cloud SQL for PostgreSQL." https://discuss.google.dev/t/optimizing-performance-and-scaling-with-managed-connection-pooling-for-cloud-sql-for-postgresql/270528
[10] GoogleCloudPlatform/cloud-sql-proxy. "Cloud SQL IAM service account authentication failed for user ..." Issue #2212. https://github.com/GoogleCloudPlatform/cloud-sql-proxy/issues/2212