Technical Post · Amazon Web Services
Agent observability on AWS: tracing the answer back to its source chunk, with the cost of the trail worked out
Someone opens a ticket saying the agent gave an answer that isn't anywhere in the knowledge base. You have the application log, you have the…
Keywords
Someone opens a ticket saying the agent gave an answer that isn't anywhere in the knowledge base. You have the application log, you have the latency, you have the token count, and you have no way to answer the only question that matters: which passages was the model reading when it wrote that. The log says the call happened. It doesn't say what the call contained.

This post builds the instrumentation that answers that question on Bedrock AgentCore, with Terraform, and works out how much it costs per month under three capture regimes.
What this post delivers
- An agent on AgentCore Runtime emitting spans to CloudWatch, with Transaction Search enabled via
awscc. - An attribute schema that links the answer to the chunk it came from without loading the chunk text into the span.
- The cost math for the three capture regimes, with us-east-1 prices checked on 08/30/2026.
- A data protection policy on the spans log group, because a user prompt is free-form input.
It does not cover quality evaluation. A trace shows what happened, not whether the answer was right. Those are two problems, and mixing them up is the most common mistake in this area.
Prerequisites
Terraform 1.9+, provider hashicorp/aws ~> 6.0 and hashicorp/awscc ~> 1.60, region us-east-1, aws-opentelemetry-distro >= 0.18.0 in the agent. Permissions to create log groups, a CloudWatch Logs resource policy, IAM roles, and the AgentCore runtime. The lab runs in under an hour, and the spend stays below one dollar if you tear everything down at the end.
The architecture
cliente
│ X-Amzn-Bedrock-AgentCore-Runtime-Session-Id
▼
┌─────────────────────────┐
│ AgentCore Runtime │ ADOT auto-instrumenta o processo
│ agente (Strands/LC) │ spans: invoke → retrieve → model → tool
└───────────┬─────────────┘
│ OTLP
▼
/aws/bedrock-agentcore/runtimes/<id>-<endpoint>
stream: spans stream: runtime-logs
│
├──► política de proteção de dados (máscara na ingestão)
│
├──► X-Ray: indexa N% como trace summary
│
└──► Application Signals / GenAI Observability
Why these components, one by one. AgentCore Runtime because it already delivers the session span and the memory span with no code, and because the session.id travels in the invocation header, which is what lets you join fifteen traces of a conversation into one investigable unit. Transaction Search because it ingests 100% of spans as structured logs and leaves X-Ray indexing as a separate decision; without it, you have X-Ray sampling and a search that can't reach the span you care about. A dedicated log group for the agent instead of the shared aws/spans, enabled by UNIFIED_TRACES_DESTINATION_ENABLED=true, because retention, data protection policy, and query scope become per-agent; with a shared destination you apply everyone's most restrictive policy to everyone, and you pay query scans over spans from agents that have nothing to do with the case. Data protection policy because the decision to capture content and the decision to store personal data are the same decision, made in the same place.
Implementation
1. Versions and tags
terraform {
required_version = ">= 1.9"
required_providers {
aws = { source = "hashicorp/aws", version = "~> 6.0" }
awscc = { source = "hashicorp/awscc", version = "~> 1.60" }
}
}
provider "aws" {
region = "us-east-1"
default_tags {
tags = {
Projeto = "observabilidade-agente"
Ambiente = "lab"
Responsavel = "cdiego"
}
}
}
2. Transaction Search
Two separate things happen here. The resource policy authorizes X-Ray to write spans to the log groups; the configuration resource defines how much of that becomes an indexed summary.
resource "aws_cloudwatch_log_resource_policy" "xray_spans" {
policy_name = "TransactionSearchXRayAccess"
policy_document = jsonencode({
Version = "2012-10-17"
Statement = [{
Sid = "TransactionSearchXRayAccess"
Effect = "Allow"
Principal = { Service = "xray.amazonaws.com" }
Action = "logs:PutLogEvents"
Resource = [
"arn:aws:logs:us-east-1:${local.account_id}:log-group:aws/spans:*",
"arn:aws:logs:us-east-1:${local.account_id}:log-group:/aws/application-signals/data:*",
"arn:aws:logs:us-east-1:${local.account_id}:log-group:/aws/bedrock-agentcore/runtimes/*",
]
Condition = {
StringEquals = { "aws:SourceAccount" = local.account_id }
ArnLike = { "aws:SourceArn" = "arn:aws:logs:us-east-1:${local.account_id}:*" }
}
}]
})
}
resource "awscc_xray_transaction_search_config" "principal" {
indexing_percentage = 5
depends_on = [aws_cloudwatch_log_resource_policy.xray_spans]
}
The wildcard in runtimes/* covers the agent log groups in this account, not all of CloudWatch. If you have agents from different teams in the same account and that is not acceptable, list the ARNs one by one and accept the friction of changing the policy with every new agent.
indexing_percentage = 5 is a choice, not a default. The service default is 1%, and that 1% is free. The cost section shows what the other 4% buys and what it costs.
3. The agent's log group, with declared retention
resource "aws_cloudwatch_log_group" "agente" {
name = "/aws/bedrock-agentcore/runtimes/${local.agente_id}-DEFAULT"
retention_in_days = 30
log_group_class = "STANDARD"
}
STANDARD class because INFREQUENT_ACCESS costs half as much on ingestion and doesn't work here: the use case is interactive querying during an incident, which is exactly what the cheap class doesn't offer. Thirty days because agent investigations happen within a short window; if your regulator asks for more, the way to go is exporting to S3, not raising the log group's retention.
4. Masking before the query exists
resource "aws_cloudwatch_log_data_protection_policy" "agente" {
log_group_name = aws_cloudwatch_log_group.agente.name
policy_document = jsonencode({
Name = "mascarar-pii-em-spans"
Version = "2021-06-01"
Statement = [
{
Sid = "Auditar"
DataIdentifier = [
"arn:aws:dataprotection::aws:data-identifier/EmailAddress",
"arn:aws:dataprotection::aws:data-identifier/CpfCode-BR",
"arn:aws:dataprotection::aws:data-identifier/PhoneNumber-BR",
]
Operation = {
Audit = { FindingsDestination = {} }
}
},
{
Sid = "Mascarar"
DataIdentifier = [
"arn:aws:dataprotection::aws:data-identifier/EmailAddress",
"arn:aws:dataprotection::aws:data-identifier/CpfCode-BR",
"arn:aws:dataprotection::aws:data-identifier/PhoneNumber-BR",
]
Operation = {
Deidentify = { MaskConfig = {} }
}
},
]
})
}
The audit statement is there so you find out what is coming in, and the masking statement is there to keep it from being readable. CloudWatch ships managed identifiers for Brazilian CPF, CNPJ, RG, CEP, and phone numbers, which covers a good share of the cases here; check the full list before relying on it. For data that has no managed identifier, such as a contract number or an internal customer identifier, the masking has to happen in code, before the attribute becomes a span.
5. The runtime
resource "awscc_bedrockagentcore_runtime" "agente" {
agent_runtime_name = "agente_rag_rastreavel"
role_arn = aws_iam_role.agente.arn
agent_runtime_artifact = {
container_configuration = {
container_uri = "${aws_ecr_repository.agente.repository_url}:${var.tag_imagem}"
}
}
network_configuration = { network_mode = "PUBLIC" }
environment_variables = {
AGENT_OBSERVABILITY_ENABLED = "true"
UNIFIED_TRACES_DESTINATION_ENABLED = "true"
OTEL_METRICS_EXPORTER = "awsemf"
RAG_INDEX_VERSION = var.versao_indice
RAG_CAPTURAR_CONTEUDO = "false"
}
}
The last two variables are mine, not the service's. RAG_INDEX_VERSION is what keeps the chunk identifier meaning something after a reindex. RAG_CAPTURAR_CONTEUDO is the switch between the regimes the next section measures, and it defaults to false for the same reason the OpenTelemetry GenAI conventions don't capture content by default.
6. The attribute that solves the problem
import hashlib, json, os
from opentelemetry import trace
tracer = trace.get_tracer("agente.rag")
VERSAO_INDICE = os.environ["RAG_INDEX_VERSION"]
CAPTURAR = os.environ.get("RAG_CAPTURAR_CONTEUDO") == "true"
def recuperar(pergunta: str, kb_id: str):
with tracer.start_as_current_span("rag.retrieve") as span:
r = cliente.retrieve(
knowledgeBaseId=kb_id,
retrievalQuery={"text": pergunta},
retrievalConfiguration={"vectorSearchConfiguration": {"numberOfResults": 5}},
)
refs = [
{
"p": i, # posição no contexto montado
"u": x["location"]["s3Location"]["uri"],
"s": round(x["score"], 4),
"h": hashlib.sha256(x["content"]["text"].encode()).hexdigest()[:12],
}
for i, x in enumerate(r["retrievalResults"])
]
# um atributo serializado, não um atributo por chunk:
# cinco atributos separados viram cinco chaves indexadas e nenhum ganho de busca
span.set_attribute("rag.chunks", json.dumps(refs, separators=(",", ":")))
span.set_attribute("rag.index_version", VERSAO_INDICE)
span.set_attribute("rag.k", 5)
if CAPTURAR:
span.set_attribute("rag.chunks_text", json.dumps([x["content"]["text"] for x in r["retrievalResults"]]))
return r, refs
This attribute comes to roughly 500 bytes for five chunks with short URIs, and about double that when the S3 paths are long. The equivalent with the text of the same five chunks comes to between 15 and 20 KB. The difference between the two is the difference between columns B and C in the table further down.
Notice what is not here: the model call span doesn't repeat the chunks. It is a child of the same trace, and the trace is the join key. Duplicating the context in every span is the silent cost multiplier in this architecture, and the next section shows how big it is.
7. IAM
data "aws_iam_policy_document" "agente" {
statement {
actions = ["bedrock:Retrieve"]
resources = [var.kb_arn]
}
statement {
actions = ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"]
resources = [var.modelo_arn]
}
statement {
actions = ["logs:CreateLogStream", "logs:PutLogEvents", "logs:DescribeLogStreams"]
resources = ["${aws_cloudwatch_log_group.agente.arn}:*"]
}
statement {
actions = ["xray:PutTraceSegments", "xray:PutTelemetryRecords"]
resources = ["*"]
}
}
The :* on the logs resource reaches the log streams inside that group, and nothing else. The X-Ray "*" is different: those two actions don't accept a resource, so the wildcard is mandatory, and it's worth knowing what it exposes, which is the ability to write trace segments across the entire account. In a multi-team account, that means a compromised agent can pollute everyone else's traces.
The proof: how much the trail back to the source costs per month
The question. An agent with one million sessions per month. How much does it cost to keep enough traceability to take a suspicious answer and get to the passage that produced it, and how much does that change if I store the text of the passages instead of their address?
The method. This is a calculation using published prices, not a measurement. us-east-1 prices checked on the CloudWatch pricing page on 08/30/2026. Stated assumptions, all replaceable with your own numbers:
- 1,000,000 sessions per month, 8 spans per session (invocation, three model calls, two retrieves, one memory operation, one gateway operation), for a total of 8,000,000 spans.
- Regime A, metadata only: 1.2 KB per span, 9.6 KB per session.
- Regime B, metadata plus chunk address (URI, score, 12-character hash, index version): 2.5 KB per span, 20 KB per session.
- Regime C, full content: 85 KB per session, accounting for the retrieved context reappearing in each of the three model calls in the same trace.
- 500 Logs Insights queries per month scanning 30 days. It sounds like a lot and it isn't: a dashboard that refreshes hourly uses 720 on its own, and an incident investigation uses 20 to 40.
- Indexing at 100% of spans, with the free 1% deducted. 30-day retention.
- For the comparison with inference at the end: 4,000 input tokens and 400 output tokens per model call, which comes to 12,000 million input tokens in the month.
The reading. Three things come out of this.
The first is that querying costs more than ingestion in every regime, and in regime C it costs seven times more. Everyone sizes observability by what goes in. The bill is dominated by what gets scanned afterward, and scanning is proportional to the size of what you stored, multiplied by the number of times someone asks. Storing the chunk text isn't expensive once. It's expensive five hundred times a month.
The second is that ingestion and indexing are independent levers. Ingestion scales with span size; indexing scales with span count. Stuffing the span with content doesn't change the indexing line by a cent, and cutting indexing from 100% to 5% takes nothing off the ingestion line. They are two knobs, and I've seen more than one team turn the wrong one.
The third is the proportion. The inference bill for this same volume is 12.000 × p_entrada + 1.200 × p_saída, with your model's per-million-token prices, which you look up on the Bedrock pricing page on the day you run the numbers. The input portion alone, at US$ 0.25 per million, already comes to US$ 3,000, and regime C is 8% of that. At US$ 3 per million, it comes to US$ 36,000, and regime C drops to 0.7%. What stops you from turning on full tracing is rarely the price, which is why the real decision in this architecture is about personal data and about querying, not about the ingestion bill.
What this number does not prove. The span sizes are assumptions, and they are the most sensitive variable in the whole model. Measure yours before trusting the table, and measuring is free: divide the IncomingBytes metric by IncomingLogEvents for the agent's log group in CloudWatch, using metric math. That gives you the real average size of your span in bytes, without scanning anything. If it comes out at 6 KB in regime B instead of 2.5 KB, the entire column triples and the conclusion about query dominance gets stronger, not weaker.
The calculation also doesn't prove that regime B answers the question in your knowledge base. It does if, and only if, the address keeps resolving. That's the subject of the next section.
Where it hurts
Reindexing erases the trail without deleting anything. You stored the chunk's URI and position. The knowledge base was reprocessed with a different window size, the boundaries shifted, and the identifier now points to different text from what the model read. The trace is still there, intact, and lying. That's why the hash and rag.index_version are in the span: the hash detects the divergence, the version explains it. Without both, regime B is theater.
The CloudWatch Logs event limit is 1,024 KB and it is not adjustable. A span with a long history plus twenty chunks goes over that. What happens is the worst possible outcome: the trace disappears precisely in the anomalous session, which is the only one you wanted to see. If you run regime C, truncate the content in code, with the size declared as an attribute, instead of discovering the limit by its absence.
The default 1% indexing is misleading. Enabling Transaction Search gives you the feeling of searching over everything. The spans are all in the log, yes. The trace summaries in X-Ray are not: only 1% of them. Trace search finds what was indexed, and the difference between “is stored” and “is searchable” only shows up in the middle of an incident.
session.id doesn't cross process boundaries on its own. If the agent calls another service, propagate it via baggage. Without that, you have two orphan traces and no session.
Head sampling discards the entire session. If you need to sample, sample per session rather than per span, and force 100% when there's a signal: an error, a triggered guardrail, negative feedback. The session you care about is always the one random sampling throws away.
Cost and operations
Regime B, indexing at 100%, 30-day retention, us-east-1, prices checked on 08/30/2026.
In the peak scenario, indexing at 5% instead of 100% takes that US$ 59.40 line down to US$ 2.40, and the Logs Insights line then accounts for 86% of a US$ 552 bill. A recurring dashboard on top of Logs Insights is the most expensive item in this architecture, and the way out is well known: a metric filter for what is always watched, a query for what is occasionally investigated.
The cost that doesn't show up on the bill: someone has to maintain the attribute schema, and an attribute schema is a contract. The day a team renames rag.chunks to retrieval.chunks, every saved query stops working, silently, and nobody notices until the next incident. This goes into the runbook or it becomes debt.
When NOT to use this
- Below 50,000 sessions per month, with a single team. The trail costs a few dollars and the pipeline costs weeks of attention. The Bedrock console and the application log will do, until the first incident that requires reconstructing an entire session. Then come back here.
- When prompt data can't reside in CloudWatch. If the retention or residency regime doesn't add up, no mask will fix it, because masking acts at ingestion and doesn't change where the data lives. Alternative: your own OTEL collector, exporting to a destination you control, accepting the loss of the ready-made integration with Application Signals.
- When the question is about quality, not about origin. A trace doesn't tell you whether the answer was right. If what you want is to measure factual correctness or relevance, the tool is an evaluation harness with a labeled dataset, and it runs outside the production path.
- When the requirement is search by arbitrary attribute over long retention. The ingestion-plus-Logs Insights pair gets expensive fast in this design, and the math above shows where. Alternative: export spans to S3 as Parquet and query them with Athena, trading query latency for price.
Wrapping up
Storing the chunk's address answers the same question as storing the chunk, for a quarter of the price, under a condition almost nobody writes down: the address has to keep resolving after the next reindex. What takes care of that is not the observability platform; it's the versioning discipline of the knowledge base.
What interests me in the math above is not the total. It's that the dominant line is querying, not collection. Agent instrumentation has been sized as if the problem were what goes in, and the problem is how many times you're going to need to ask.
See you in the next post!
References:
- Understand observability for agentic resources in AgentCore
- Add observability to your Amazon Bedrock AgentCore resources
- CloudWatch Transaction Search and Enable Transaction Search
- Amazon CloudWatch Pricing (checked on 08/30/2026)
- CloudWatch Logs quotas
- Generative AI observability now generally available for Amazon CloudWatch
- Inside the LLM Call: GenAI Observability with OpenTelemetry
- awscc_xray_transaction_search_config and awscc_bedrockagentcore_runtime
Comments
Every comment is moderated before it appears here. Nothing is published automatically.
Loading…