← All posts

Technical Post · Amazon Web Services

Amazon Bedrock AgentCore: the operational “core” for agents in production

If you have ever run an agent locally (LangGraph, Strands, CrewAI, etc.), you know: the problem is rarely “getting the model to…

8 min read1,694 wordsSections: 5Images: 1Code blocks: 9Feb 7, 2026

Keywords

Share
Comment

If you have ever run an agent locally (LangGraph, Strands, CrewAI, etc.), you know: the problem is rarely “getting the model to respond.” The problem is operating it with security, traceability, isolation, access control, versioning, stable endpoints, observability, and governance, without it turning into a festival of scripts and exceptions.

Amazon Bedrock AgentCore exists precisely to remove that layer of “undifferentiated heavy lifting” and give you composable services for taking agents to production. At its heart is AgentCore Runtime, which hosts your agent as a serverless workload with per-session microVM isolation, support for long-running executions, authentication, and agent-oriented observability trails.

Below I get straight to the point: how to design the foundation (network, IAM, keys, logs, versioning, and pipelines) and how to provision runtimes and endpoints in Terraform without staying on the surface.

1) What you are actually provisioning

1.1 Runtime ≠ “the agent”

In AgentCore, the Runtime is your agent's execution environment (or that of a tool / MCP server). You “drop off” an artifact (a container in ECR or code via S3), and the Runtime then exposes an invocation contract.

The Runtime comes with properties that matter operationally:

  • NetworkConfiguration: PUBLIC or VPC (yes, this completely changes your threat model).
  • LifecycleConfiguration: per-session time limits and maximum lifetime. (Essential for cost, resilience, and security.)
  • EnvironmentVariables: controllable, but they require discipline (don't let them become an accidental “secrets vault”).
  • AuthorizerConfiguration: inbound authorization options.

1.2 Versioning and endpoints (what almost everyone neglects)

Every time you create a Runtime, you get version 1. Every update produces a new immutable version. Endpoints are pointers to versions — including a DEFAULT one that points to the latest version.

This is the basis of your promotion model:

  • dev → staging → prod
  • the prod endpoint should not “float” (unless you want it to), while DEFAULT can be used for fast iteration.

1.3 Invocation: Runtime ARN + payload (and streaming support)

Invocation uses InvokeAgentRuntime, with support for sessions and real-time/streaming responses.

2) The minimum foundation for provisioning agents “the right way” (no fluff)

I will organize the foundation in layers, because that is how you avoid rework as the number of agents grows.

Layer A — Account/environment standards and governance

Goal: predictable environments, separated blast radius, and traceability.

  • Separate by account (at minimum: shared, dev, prod). If that is not possible, separate by workspaces + strict naming/tagging.
  • Mandatory tags: app, env, owner, cost_center, data_classification, criticality.
  • KMS per domain (logs/secrets/sensitive data) and least-privilege policies.

Layer B — Network and egress (where “PUBLIC” turns into debt)

Critical decision: PUBLIC vs VPC.

For enterprise, the default tends to be VPC in order to:

  • control egress,
  • use private endpoints,
  • reduce the exposure surface.

The Runtime itself supports NetworkMode: PUBLIC | VPC.

Practical recommendation:

  • Private subnets + NAT with controlled egress (or egress through a firewall).
  • VPC Endpoints (Interface/Gateway) for whatever your agent uses (S3, ECR, Logs, STS, Bedrock as applicable).
  • “Deny by default” SGs, opening only what is needed.

Layer C — IAM (where agents silently break security)

At a minimum, you will have:

  1. Runtime execution role (the runtime's roleArn).
  2. Model invocation permissions (bedrock:InvokeModel* when the agent calls models on Bedrock). This even shows up in updates to AgentCore-related managed policies.
  3. Tool-specific permissions (S3, DynamoDB, internal APIs, etc.) — always in a separate policy per “capability,” not one giant monster.

Operational rule: the “runtime role” should not have permission to administer infrastructure. It executes. Period.

Layer D — Agent observability (it's not just logs)

AgentCore Runtime provides trails focused on “agent reasoning / tool invocations / model interactions” for debugging and auditing.

Recommended foundation:

  • Log groups standardized by env/app/runtime.
  • Retention by class (dev: short, prod: aligned with compliance).
  • Correlation by sessionId, requestId, traceId (OpenTelemetry where applicable).

Layer E — Versioning, endpoints, and promotion (CI/CD)

If you don't automate this, it becomes “artisanal deployment”:

  • Build the artifact (container) → push to ECR
  • Update the Runtime (produces a new version)
  • Update the staging endpoint
  • Smoke test (invoke)
  • Manual/approved promotion → update the prod endpoint

The endpoints/versions mechanism exists precisely for this.

3) Provisioning in Terraform (with Runtime + Endpoint + foundation)

3.1 Why I am using awscc here

At the moment, the AgentCore Runtime/Endpoint resources show up naturally through CloudFormation, and in the Terraform ecosystem the most direct path is the awscc provider (Cloud Control / CloudFormation-backed). That gives you a stable track while the “main” aws provider evolves to cover everything natively.

_If you would rather not depend on awscc, the pragmatic alternative is to trigger creation via the AWS CLI (bedrock-agentcore-control create-agent-runtime etc.) in a null_resource — but then you are handling drift and idempotency by hand._

/infra
  /modules
    /agentcore-foundation
    /agentcore-runtime
  /envs
    /dev
    /prod
  • agentcore-foundation: VPC, endpoints, KMS, ECR, logs, base IAM
  • agentcore-runtime: runtime + endpoints + agent-specific policies

3.3 Code: providers, conventions, and tags

terraform {
  required_version = ">= 1.6.0"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = ">= 5.0"
    }
    awscc = {
      source  = "hashicorp/awscc"
      version = ">= 0.80.0"
    }
  }
}

provider "aws" {
  region = var.region
}

provider "awscc" {
  region = var.region
}

locals {
  tags = {
    app                = var.app_name
    env                = var.env
    owner              = var.owner
    cost_center        = var.cost_center
    data_classification = var.data_classification
  }
}

3.4 Foundation: ECR (runtime artifact) + KMS + Logs

ECR (with immutability and scanning)

resource "aws_ecr_repository" "agent" {
  name                 = "${var.app_name}-${var.env}-agentcore"
  image_tag_mutability = "IMMUTABLE"

  image_scanning_configuration {
    scan_on_push = true
  }

  tags = local.tags
}

KMS (for encrypting logs/secrets/data)

resource "aws_kms_key" "agentcore" {
  description             = "KMS for AgentCore (${var.app_name}/${var.env})"
  deletion_window_in_days = 30
  enable_key_rotation     = true
  tags                    = local.tags
}

CloudWatch Logs (retention standard)

resource "aws_cloudwatch_log_group" "agentcore" {
  name              = "/agentcore/${var.env}/${var.app_name}"
  retention_in_days = var.log_retention_days
  kms_key_id        = aws_kms_key.agentcore.arn
  tags              = local.tags
}

Note: AgentCore may create its own logs in specific namespaces (including for evaluations). If you want “everything under control,” plan naming/retention and permissions explicitly.

3.5 IAM: Runtime Role (real least privilege)

data "aws_iam_policy_document" "runtime_assume" {
  statement {
    effect = "Allow"
    actions = ["sts:AssumeRole"]

    principals {
      type        = "Service"
      identifiers = ["bedrock-agentcore.amazonaws.com"]
    }
  }
}

resource "aws_iam_role" "runtime" {
  name               = "${var.app_name}-${var.env}-agentcore-runtime"
  assume_role_policy = data.aws_iam_policy_document.runtime_assume.json
  tags               = local.tags
}

# Exemplo: permitir invocar modelos no Bedrock
data "aws_iam_policy_document" "runtime_permissions" {
  statement {
    sid     = "BedrockInvoke"
    effect  = "Allow"
    actions = ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"]
    resources = var.allowed_model_arns
  }

  # Exemplo: logs
  statement {
    sid    = "WriteLogs"
    effect = "Allow"
    actions = [
      "logs:CreateLogStream",
      "logs:PutLogEvents"
    ]
    resources = ["${aws_cloudwatch_log_group.agentcore.arn}:*"]
  }
}

resource "aws_iam_policy" "runtime_permissions" {
  name   = "${var.app_name}-${var.env}-agentcore-runtime-permissions"
  policy = data.aws_iam_policy_document.runtime_permissions.json
  tags   = local.tags
}

resource "aws_iam_role_policy_attachment" "runtime_attach" {
  role       = aws_iam_role.runtime.name
  policy_arn = aws_iam_policy.runtime_permissions.arn
}

3.6 Provisioning the Runtime (awscc / CloudFormation-backed)

CloudFormation specifies that the Runtime requires:

  • AgentRuntimeName
  • AgentRuntimeArtifact (container or code in S3)
  • RoleArn
  • NetworkConfiguration (PUBLIC or VPC)
  • (optional) LifecycleConfiguration, EnvironmentVariables, AuthorizerConfiguration, etc.

Example with a container in ECR and VPC mode

resource "awscc_bedrockagentcore_runtime" "this" {
  agent_runtime_name = "${var.app_name}_${var.env}"

  description = "AgentCore Runtime for ${var.app_name} (${var.env})"

  role_arn = aws_iam_role.runtime.arn

  # Artefato: container
  agent_runtime_artifact = {
    container_configuration = {
      container_uri = "${aws_ecr_repository.agent.repository_url}:${var.image_tag}"
    }
  }

  # Rede: VPC
  network_configuration = {
    network_mode = "VPC"
    network_mode_config = {
      vpc_id     = var.vpc_id
      subnet_ids = var.private_subnet_ids
      security_group_ids = [var.runtime_sg_id]
    }
  }

  # Ciclo de vida (controle de custo / segurança)
  lifecycle_configuration = {
    idle_runtime_session_timeout = var.idle_session_timeout_seconds
    max_lifetime                 = var.max_session_lifetime_seconds
  }

  environment_variables = {
    ENV        = var.env
    LOG_LEVEL  = var.log_level
  }

  tags = local.tags
}

The AgentCore Runtime API explicitly exposes networkConfiguration, lifecycleConfiguration, environmentVariables, and roleArn as creation/control parameters.

3.7 Stable endpoint for promotion (dev/staging/prod)

The RuntimeEndpoint resource in CloudFormation takes:

  • AgentRuntimeId
  • Name
  • (optional) AgentRuntimeVersion and Description

In Terraform (awscc), it looks like this:

resource "awscc_bedrockagentcore_runtime_endpoint" "prod" {
  agent_runtime_id = awscc_bedrockagentcore_runtime.this.agent_runtime_id
  name             = "prod"
  description      = "Stable production endpoint for ${var.app_name}"

  # Fixar versão (recomendado em prod)
  agent_runtime_version = awscc_bedrockagentcore_runtime.this.agent_runtime_version

  tags = local.tags
}

How I run this in practice (no drama)

  • In dev, I can let the endpoint “float” (point to latest) or use DEFAULT.
  • In prod, I pin the version and only update the endpoint after the pipeline and approval.

This model is exactly what AgentCore describes: endpoints as pointers, immutable versions, and DEFAULT tracking the latest.

3.8 Invocation: automated smoke test in the pipeline

You can test after provisioning using the AWS CLI (or SDK). The CLI exists for create-agent-runtime and also for invoking.

Invoke example (JSON payload):

aws bedrock-agentcore invoke-agent-runtime \
  --agent-runtime-arn "$RUNTIME_ARN" \
  --payload fileb://payload.json \
  --content-type application/json \
  /dev/stdout

The InvokeAgentRuntime API is the contract for real-time requests/responses and supports a qualifier for the endpoint/version.

4) Practical checklists: what separates “provisioning” from “provisioning right”

4.1 Security (baseline)

  • Runtime in a VPC whenever there is sensitive data or internal integration (the enterprise default).
  • “Capability-based” IAM: small policies per resource (S3 read, Dynamo read, etc.).
  • Secrets outside env vars (an env var is config, not a vault).
  • Auditing: logs/traces with appropriate retention and per-session correlation.

4.2 Operations

  • prod endpoint pinned to a version.
  • Deploy = new immutable container tag + runtime update + endpoint update.
  • Explicit timeouts (idle and max lifetime) to avoid “zombie” sessions.

4.3 Quality (before scaling)

AWS has published a very down-to-earth set of best practices for enterprise agents — instrument from day 1, a deliberate tool strategy, automated evaluation, combining the agent with deterministic code, etc. I recommend using it as a maturity checklist for when you go from 1 agent to an “agent platform.”

5) Wrapping up: the central idea

AgentCore is not “yet another way to write prompts.” It is an operational layer for agents with:

  • a runtime with per-session isolation and long-running execution,
  • automatic versioning + endpoints,
  • and an infrastructure surface (network/IAM/observability) that you need to treat as an internal product.

See you in the next post!

Comments

Every comment is moderated before it appears here. Nothing is published automatically.

Loading…